Interestana
Home/News/Perplexity Releases pplx-embed-v2-context-9b-preview Embedding Model
MarkTechPost••3 min read

By Interestana AI Editorial — AI-drafted, human-overseen. How we report

Perplexity Releases pplx-embed-v2-context-9b-preview Embedding Model

Perplexity Research, in collaboration with turbopuffer, has released pplx-embed-v2-context-9b-preview, a new contextual embedding model designed for Retrieval Augmented Generation (RAG) pipelines. This model distinguishes itself by embedding each chunk of text with the full document in view, a departure from traditional methods. The core innovation lies in its training signal, which teaches the model to retrieve not just a relevant passage, but the specific context necessary to verify an answer. This approach aims to overcome limitations of RAG systems that rely on identifying a single 'gold passage' for each query.

The pplx-embed-v2-context-9b-preview model is currently available as a self-hosted preview, with its weights accessible on Hugging Face under the MIT license. To load the model, users require transformers version 5.4.0 or later and must set trust_remote_code=True. It is not yet integrated into the Perplexity API, and the model card indicates that weights and the interface may undergo changes without prior backward compatibility.

Traditional RAG systems typically split long documents into smaller chunks for processing. However, the relevance of a chunk often depends on information presented elsewhere in the document, such as entities, headings, or definitions. Contextual models, like this new release, address this by employing a 'late chunking' strategy. In this method, the entire document is encoded in a single pass, and then embeddings are pooled for each chunk. The training process for such models usually involves marking one 'gold chunk' per query, treating all other chunks, including those that provide verifiable context, as negative examples. Perplexity identifies three key issues with this 'gold passage' approach: binary labels offer a coarse training signal, the cost of LLM annotation scales linearly with dataset size, and labels become intrinsically tied to a specific chunking strategy.

The training methodology for pplx-embed-v2-context-9b-preview utilizes Perplexity’s query-aware context compression model as a teacher. This teacher model processes the query and document together, assigning a score to each token. Chunk relevance is then calculated as the mean of the top 'n' token scores within each chunk. A 'soft target' is generated using a temperature-scaled softmax over the chunks within the positive document, with chunks from other documents receiving zero scores. The training incorporates a distillation loss, measured by the forward KL divergence between the teacher and student distributions, and a document loss using InfoNCE, where a document is scored based on its best chunk, drawing inspiration from ColBERT's MaxSim. Each training batch employs a random chunking strategy. Chunks are demarcated by a learned '<|chunk_sep|>' token and processed via mean-pooling. Crucially, the teacher model operates exclusively during training, meaning its use does not introduce any additional latency or storage requirements during inference. The model itself is initialized from an in-house 9B ColBERT model.

Original source — read the full reporting at the publisher:

Read on MarkTechPost

Get the weekly AI digest

AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.

Read next