AI News
AI News AgentModel releaseGoogle3 min read

Google launches EmbeddingGemma 2, a local multimodal AI model

Google introduces EmbeddingGemma 2, an open model with 740 million parameters that searches and relates text, code, images, video and audio directly on local devices. It promises multimodal searches without the cloud, lower memory usage and an 8,000-token context window.

Google has launched EmbeddingGemma 2, an open AI model that lets you search and relate text, code, images, video and audio directly on devices such as phones and computers.

The model converts each type of content into a numerical representation known as an embedding. When two pieces of content have a similar meaning, their representations sit close to each other. This lets you search for a specific video by typing a sentence, find a voice recording from a text query or locate related code within a project.

A model designed to run without the cloud

EmbeddingGemma 2 has 740 million parameters and is released under the Apache 2.0 license, which allows you to use and adapt it in commercial projects as well. It is based on the Gemma 4 architecture and was designed to run on local hardware.

That has three practical effects:

  • Your data can stay on the device, which matters for recordings, private documents or internal code.
  • Searches do not necessarily depend on an internet connection.
  • Responses can be faster because each file does not have to be sent to a server.

Google says the first version of EmbeddingGemma surpassed 20 million downloads. The new version expands its scope by combining multiple formats in the same search space.

Less memory and more formats

The model is modular. For text-only tasks, it can use a configuration with 270 million parameters, while the optional vision and audio components add 170 and 300 million, respectively.

It also uses a technique called Matryoshka Representation Learning, which lets you reduce the size of output vectors without processing the content again. Instead of storing 768 dimensions, developers can choose 512, 256 or 128. According to Google, this can reduce the storage and memory used by a local vector database by as much as six times.

On a Google Pixel 11 Pro, with quantization, the model requires approximately 191 MB of active RAM in its text version and around 567 MB in the full multimodal version.

More context for audio, images and video

The context window reaches 8,000 tokens, four times more than in EmbeddingGemma. Under compatible conditions, it can process up to:

  • 5.5 minutes of audio.
  • 29 images.
  • 58 video frames.
  • Combinations of these formats.

The improvement is also visible in code. In the MTEB Code test, Google says the score rises from 68.76 to 78.68, an improvement of 9.92 points over the previous generation. This can help index repositories, perform semantic searches and give programming agents more context.

What you can build with it

One use case is an application that finds a scene within hours of video based on a voice note. Another is a local search tool that relates manuals, screenshots, recordings and code snippets without uploading them to the cloud.

Combined with a generative model such as Gemma 4, it can also form local RAG systems. These systems first retrieve relevant information from your files and then provide it to a model so it can draft a response.

The weights are available on Hugging Face and Kaggle. Google also indicates compatibility with tools such as MediaPipe, LiteRT, transformers.js, WebGPU, llama.cpp, Ollama and LM Studio.

What matters is not only that the model understands more formats, but that it can do so on relatively modest hardware. The next step will be to see whether real-world applications can maintain the quality Google reports in its tests while combining privacy, multimodal search and low resource usage.