Perplexity Open-Sources pplx-embed-v2-late, 92.4% on MADQA
TL;DR
- Both models share a 128-dim-per-token space distilled from an 18B ColBERT teacher; the 0.6B queries 9B-built indexes without rebuilding the index.
- MADQA's 92.4% headline obscures a drop to 61-65% nDCG@10 on markdown documents, a gap that matters for structured-data pipelines.
- Token-level ColBERT vectors multiply index storage over single-vector models; adoption economics depend on whether query volume or data volume drives cost.
Perplexity has released two open-weight embedders, a 0.6B edge model and a 9B, that share one embedding space, so an index built with the big model can be served by the small one. Both landed October 7 on Hugging Face under the MIT license, announced on the company forum.
The pitch is retrieval over text, images, and rendered pages without a chunking or OCR pipeline in front. Perplexity's model card describes them as "late-interaction (ColBERT) retrievers for text, images, and visual documents, built on Qwen3.5 with bidirectional attention," each producing "one 128-dimensional vector per token" scored with MaxSim.
The 9B scored 92.4% on MADQA, a set of 500 questions over roughly 18,000 PDF pages, and 64.0% on BrowseComp+ paired with GPT-OSS-120B. AlphaSignal's writeup puts that 4.9 points ahead of competing ColBERT models. On Q2D-Web the company lists 74.8% Recall@1000 across about 190 million documents.
Both weights come from a bigger teacher. "Both models were distilled from an internal 18B ColBERT teacher trained on pair and triplet data. Distillation used a token-level LEAF-style objective," the model card says. The 0.6B was fully fine-tuned; for the 9B, the final eight transformer layers were fully fine-tuned while the remaining layers and the vision encoder were adapted with LoRA.
The 0.6B lands at 594M total parameters, with 240M for text and 340M for images. The 9B card lists 7.4B active parameters out of 8B. API access is announced as forthcoming; today the release is for self-hosters.
It arrives a day after Google's EmbeddingGemma 2 and alongside a run of retrieval and RAG work we have been tracking this week, including an Aalto result on inverting ColPali indexes.
What others are reporting
-
Hugging Face Read →
First-party model card detailing LEAF-style distillation from an 18B teacher, LoRA on the 9B's earlier layers, and per-benchmark nDCG@10 splits.
The models produce one 128-dimensional vector per token and score query-document similarity using MaxSim.
-
Crypto Briefing Read →
Frames cross-model querying as a direct cloud cost argument and flags benchmark verification and storage cost as watchpoints for production adoption.
you can index once with the heavyweight model, then answer live queries with the smaller, faster one.
-
TILNOTE Read →
Surfaces integration friction: sentence-transformers 6.0.0+ hard dependency, markdown document performance drop to 61-65% nDCG@10, and a storage-vs-latency decision framework.
Two models share one embedding space, enabling lightweight 0.6B edge devices to query indexes built by the more powerful 9B model without accuracy loss.
Originally reported by perplexity.ai
Read the original article →Original headline: Perplexity Open-Sources pplx-embed-v2-late, 0.6B and 9B Multimodal Late-Interaction Embedders