By Interestana AI Editorial — AI-drafted, human-overseen. How we report
Perplexity AI Releases New Embedding Models
Perplexity AI released pplx-embed-v2-late, a new suite of ColBERT-style multimodal embedding models, on an unspecified date, offering two distinct sizes: a 0.6 billion parameter model optimized for speed and cost-efficiency, and a 9 billion parameter model designed for maximum quality. Both models are capable of retrieving text, images, and rendered PDF pages, and importantly, they share a unified embedding space, allowing for cross-modal understanding. These models are available for self-hosting under the MIT license, which permits commercial use, and are accessible on Hugging Face. Perplexity has indicated plans for a hosted API endpoint, though it is not yet live.
The 0.6B model is engineered to function as a lightweight query encoder, suitable for deployment on edge devices or laptops, utilizing approximately 340 million active parameters for image processing. This smaller model aims to maintain performance close to larger, 8 billion parameter rivals. A key feature highlighted is the ability to search a 9B index using 0.6B queries, which recovers about half of the quality gap observed in text retrieval at the reduced query cost. The models generate 128-dimensional token vectors, significantly narrower than competitors' vectors which range from 2,048 to 4,096 dimensions, representing a 16x to 32x reduction in dimensionality.
However, the models present certain limitations. The architecture stores one vector per token, leading to index sizes that scale directly with document length. In terms of performance benchmarks, pplx-embed-v2-late is not the top performer on the ViDoRe v3 image retrieval task; Tencent's EVIE model reportedly scores higher. Furthermore, a single input cannot simultaneously contain both text and images. Perplexity has stated that all performance scores are self-reported, and the accompanying technical report has not yet been published.
Technical specifications reveal that the pplx-embed-v2-late-0.6b model comprises 594 million total parameters, with approximately 240 million active parameters for text and 340 million for images. The base model is derived from Qwen3.5-0.8B, pruned to 12 text layers. The 9B variant has 9 billion total parameters (though Hugging Face lists it as 8B) and utilizes 7.4 billion parameters. Both output 128 dimensions per token. Estimated memory requirements for weights in bf16 precision are around 1.2 GB for the 0.6B model and 16 to 18 GB for the 9B model. The models require sentence-transformers version 6.0.0 or later and transformers version 5.4.0 or later. The published checkpoints are distributed in F32 format, which doubles their download size.
Original source — read the full reporting at the publisher:
Read on MarkTechPostGet the weekly AI digest
AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.