Perplexity improves search with contextual embeddings
Perplexity introduces an embedding model that retrieves both an answer and the context needed to verify it. Its preview performs better on a long-document benchmark and supports `int8` vectors to reduce storage requirements.

Perplexity has introduced pplx-embed-v2-context-9b-preview, a model designed to help search systems find not only the sentence containing an answer, but also the context needed to understand and verify it.
The idea addresses a common problem in internal search engines and AI systems that query documents. A sentence may answer a question but be useless on its own: it might mention “the company” without saying which one, use a pronoun whose referent appears several pages earlier, or belong to a table whose header appears in another section.
The problem with searching by fragments
Long documents are often split into small fragments so they can be indexed and searched quickly. Each fragment is converted into a vector, a numerical representation of its meaning, and the system compares those vectors with the user’s query.
The problem is that the fragment loses some of its original context. If a sentence says “coverage lasts 18 months,” you will not necessarily know whether it refers to residential or commercial installations, or to a specific product. Storing the entire document as a single vector does not solve the problem either: one representation cannot accurately reflect all its details.
Contextual embeddings attempt to fill that gap. They represent each fragment while taking the full document into account, even though the system continues to store and retrieve individual fragments afterward.
From a single answer to the evidence you need
Most models are trained with a simple method: for each question, one correct fragment, known as the gold passage, is identified. The model learns to rank it above the others.
That approach treats all the remaining fragments as irrelevant, even when one of them provides information that is essential for interpreting the answer. In a contract, for example, the rental amount may appear in one sentence, while the address, tenant’s name and expiration date appear elsewhere in the document.
Perplexity’s new model learns to retrieve both types of content:
- The fragment containing the answer.
- The fragments that clarify which entity, date or version it refers to.
- The evidence needed to verify that the answer is correct.
To train it, Perplexity uses a context compression model as a teacher. That model assigns a relevance score to each word in the document based on the query. Those scores are then combined to teach the embedding which fragments are most important and which provide only partial support.
The teacher is used only during training. During search, the model produces one embedding per fragment without adding a compression or reranking stage, so it does not increase inference cost. It also offers 1,024- and 2,048-dimensional vectors, as well as an int8 quantized version that reduces the storage space they require.
Results on a context-focused benchmark
Perplexity evaluated the model on context-bench, a benchmark created and privately managed by turbopuffer. It includes 2,099 queries across 38,894 long documents, with 2,458,072 one-sentence fragments covering 21 areas, including legal contracts, clinical research, technical documentation and corporate archives.
The benchmark measures three different capabilities:
- Finding the correct document among nearly identical documents.
- Retrieving the fragment that contains the answer.
- Retrieving the evidence needed to interpret or verify it.
According to results published by Perplexity, the model achieved 45.5% answer retrieval, 40.6% evidence retrieval and 31.1% complete evidence retrieval on context-bench when the top 10 results were reviewed. It also scored 61.6% on placing the correct document among those 10 results.
The company says the model outperformed voyage-context-4 by 14.4 percentage points on answer retrieval and by 5 points on evidence retrieval at the same cutoff. The evaluation was a blind submission, and the benchmark was not used during development, according to Perplexity. Because the dataset remains private, other teams must ask turbopuffer to evaluate their models.
What changes for you
The improvement matters most in systems that answer questions using your own documents: enterprise search engines, legal assistants, contract analysis tools and systems that query technical manuals.
Instead of returning an isolated sentence such as “the deadline is 45 days,” the search engine is more likely to also retrieve the paragraph explaining that the deadline applies after a specific amendment. That reduces ambiguous answers and makes it easier for someone to quickly check whether the AI found the correct document.
The model also offers a practical storage advantage. Perplexity says its 1,024-dimensional int8 vectors take up 1 KB per vector, compared with 8 KB for a 2,048-dimensional float32 vector from another system used for comparison. The figure refers to the vector, not the complete index.
The model is available as a preview on Hugging Face. Perplexity is also working to add it to its API, but has not yet presented that availability as a finished launch.
The direction indicated by this news is clear: retrieval systems are beginning to assess not only which fragment contains the answer, but also what information makes that answer meaningful. For AI assistants, finding the data will be only half the job. The other half will be showing that it belongs to the correct document and can be supported by its context.