AI News
AI News AgentModel releasePerplexity5 min read

Perplexity launches multimodal embeddings for AI

Perplexity introduces `pplx-embed-v2-late`, a family of AI models that searches text, images, and PDF pages using multiple vectors per document. Its 0.6B and 9B versions can be combined to preserve index quality while reducing query costs.

Perplexity has introduced pplx-embed-v2-late, a new family of AI models that can search text and images without reducing each document to a single vector. The technology is designed to find answers in web pages, PDFs, screenshots, tables, and charts with greater precision.

The models are publicly available on Hugging Face, in two sizes: 0.6B and 9B parameters. The smaller model prioritizes speed and can run even on devices with limited resources. The larger one delivers higher quality when computing cost is not the main concern.

What changes compared with traditional search

AI-based search systems usually convert each document into a vector, a list of numbers that summarizes its meaning. They then compare that vector with the user’s query vector to decide which results are relevant.

The problem is that a long document can contain many topics and details. Compressing everything into a single representation can cause important connections to disappear. Splitting the document into chunks helps, but it can also separate information that only makes sense when viewed together.

pplx-embed-v2-late keeps multiple vectors for each document, one for every relevant token or text unit. It then compares each part of the query with the document’s parts and adds up the best matches.

For example, a query such as “revenue from the European division in 2024” could relate to a specific figure in a table, a section heading, and a footnote. A single vector might blend those elements together. Perplexity’s approach tries to keep them distinct.

It also understands visual pages

The other major difference is that the models work with text and images within the same search space. This lets you enter a query and directly find a rendered page from a PDF without first converting all its content into text with OCR.

OCR is the system that recognizes letters in images or scanned documents. Although useful, it can make mistakes and often loses some visual information, such as:

  • The position of a figure within a table.
  • The relationship between a chart and its explanation.
  • The layout of a page or slide.
  • The spatial structure of forms and scanned documents.

In practice, this could help you search for a figure in a financial report, locate a specific slide in a presentation, or find a scanned page even when its text was not recognized correctly.

The models can also search natural images, such as photographs, although Perplexity says their main use will be searching visual documents.

A large model for indexing and a small one for queries

One of the most useful features is that the 0.6B and 9B versions share the same embedding space. This means a company can create its document index with the large model and run daily searches with the small one.

An index is the structure that organizes documents so they can be found quickly. Creating that index usually happens once, while encoding queries happens every time someone performs a search.

A practical setup would therefore be:

  • Use pplx-embed-v2-late-9b to process documents once.
  • Use pplx-embed-v2-late-0.6b to convert each query in real time.
  • Keep the quality of the large model’s index without paying the large model’s cost for every search.

In Perplexity’s tests, this setup improved the average across 72 specialized search tasks by 1.6 percentage points compared with using the small model for everything. On the ViDoRe V3 visual benchmark, it rose from 62.3% to 63.5%.

What the tests show

Perplexity evaluated the models on text search, visual documents, images, and tasks where an AI agent needs to retrieve information before responding.

On its Q2D-Web web benchmark, based on around 190 million documents and nearly 70,000 queries, the models achieved a retrieval score of 74.8% for 9B and 73.6% for 0.6B, compared with 69.3% for the previous highest result they evaluated.

On ViDoRe V3, which focuses on finding information within visual pages, the 9B model scored 65.2% and the 0.6B model 62.3%. Perplexity especially highlights the smaller model’s result: it activated around 340 million parameters when processing images, despite competing with much larger systems.

On complex document questions, the 9B model reached 92.4% accuracy on MADQA, a test with 800 PDFs and more than 18,000 pages. The 0.6B model scored 90.1%.

The trade-off: higher quality, more storage

Multi-vector embeddings are more expressive, but they are also more expensive to store and compare. A traditional system stores one vector per document. This approach stores vectors for many tokens, so the cost grows with the length of the content.

That means it is not an automatic replacement for every search engine. For huge corpora and very fast queries, single-vector embeddings may still be simpler and cheaper. The new models make more sense when details, context, and visual information matter.

The two models were trained on 186 million query and document pairs, sourced from 594 datasets in 46 languages. Perplexity says it excluded data associated with the benchmarks used to evaluate them, with the goal of preventing the tests from artificially favoring its own models.

For you, the most visible change will not be in a search interface, but in what it can find underneath: data contained in a table, an image, or a page layout that could previously be lost when the document was converted into text. The next step will be to see how these models perform when Perplexity integrates them into its API and large-scale search systems.

Perplexity launches multimodal embeddings for AI | neversleep.ai