huggingface.co web signal

Aalto Paper Inverts ColPali Indexes, Re-IDs 98.4% of Pages

Cybersecurity RAG ai-security

TL;DR

  • On ViDoRe v3, a flow-matching inverter trained against raw ColPali-style indexes recovered 47% of words and 45% of sensitive tokens from stored vectors alone.
  • Reconstructed pages used as queries against the same stored indexes ranked their source page first 98.4% of the time, and 70.2% against a different multi-vector retriever.
  • Token pooling and shuffling both cut word recall to about 8%, but a position model raised shuffled-index re-identification back from 3.8% to 93.5%.

Stored vector indexes are not safely abstract. On the ViDoRe v3 benchmark, pages reconstructed from raw ColPali-style indexes alone recovered 47% of the words and 45% of the sensitive tokens (numbers, acronyms, capitalized terms); used as queries against the same stored indexes, those reconstructions ranked their source page first 98.4% of the time.

"Multi-vector visual document retrievers are therefore vulnerable to inversion through their stored index, which should be protected like the documents it encodes," write Yao Zhang and Yu Xiao of Aalto University in the paper.

The attack frames inversion as conditional document image generation. A flow-matching inverter learns to render a page from the roughly one thousand patch vectors a multi-vector retriever stores per document, inferring from the index itself the encoder, the page shape, and — for shuffled vectors — their order.

Two "cheap protections" are tested. Token pooling at factors of three and nine both cut word recall to about 8%. Shuffling, on its own, drops re-identification to 3.8%. But a position model that reorders a shuffled index raises that back to 93.5%, "while inverting a pooled index remains open." Applied unchanged to a second multi-vector retriever, ColQwen3.5-4.5B, the attack still ranks the source page first 70.2% of the time, though its word recall stays below a nearest-neighbor baseline.

The operational threat is hosted vector databases with separated access tiers, and the authors address remediation to store operators rather than to model vendors. The paper lists noise injection, quantization, keyed random projections, and encrypted retrieval as defenses it did not evaluate. It lands in a run of multi-vector-retrieval work our RAG tracker has been following through 11 stories in the last 90 days.