Perplexity AI Releases pplx-embed-v2-late: A 0.6B Edge Model and a 9B Model Scoring 92.4% on MADQA
Perplexity drops dual-size embeddings: 0.6B for edge, 9B scoring 92.4% on MADQA—both MIT-licensed and self-hostable.

Why it matters
Open-weight embedding models at scale compete with proprietary alternatives. Practitioners can now ground retrieval systems locally without vendor lock-in, though ViDoRe performance (61.2%) reveals task-specific limits.
The key facts
10 to knowpplx-embed-v2-late: two sizes (0.6B edge, 9B production)
MADQA benchmark: 92.4% (best score)
ViDoRe v3 Markdown: 61.2% (weakest score)
MIT license, self-hostable
No pricing, no API gate mentioned
Two model sizes: 0.6B (edge devices) and 9B (high-quality indexes)
Best benchmark score: 92.4% on MADQA
Weakest score: 61.2% on ViDoRe v3 Markdown
MIT license; self-hostable
No pricing, no API tier, no regional availability limits disclosed
The story so far
Earlier coverage of this storyline
- Google DeepMind Releases EmbeddingGemma 2, a 740M Open Multimodal Embedding Model Built on Gemma 4MarkTechPost
- EmbeddingGemma 2: an open, lightweight multimodal embedding modelGoogle DeepMind Blog
- EmbeddingGemma 2Simon Willison
- Google expands EmbeddingGemma beyond text to images, audio and videoSiliconAngle
- Google DeepMind Launches EmbeddingGemma 2 for On-Device Multimodal SearchTechRepublic
- This story
Go to the source
MarkTechPostmarktechpost.com
Publisher excerpt: Perplexity's pplx-embed-v2-late comes in 2 sizes: a 0.6B model built to run on edge devices, and a 9B model for building high-quality indexes. Its best score is 92.4% on MADQA, and its weakest is 61.2% on ViDoRe v3 Markdown. Both are MIT-licensed and ready to self-host.