EmbeddingGemma 2 is Google DeepMind's new open model for local semantic search. In a post dated October 6, 2026, research engineers Sahil Dua and Henrique Schechter Vera describe a 740-million-parameter model under the Apache 2.0 license that places text, code, images, audio and video in the same vector space.
What happened?
Embeddings are lists of numbers that represent the meaning of a piece of content. Similar items sit close together, which lets an app find a video clip from a sentence or an audio file from an image without a server. The official post is on the Google blog, and the technical docs are at ai.google.dev/gemma/docs/embeddinggemma.
The previous EmbeddingGemma was text-only and passed 20 million downloads, Google says. Version 2 is built on the Gemma 4 architecture and extends context to 8,192 tokens, four times the first model. Google says that window covers up to 5.5 minutes of audio, 29 images or 58 video frames in one call.
Why it matters
The EmbeddingGemma 2 design is modular. The text core has 270 million parameters. Vision and audio are optional encoders, at 170 million and 300 million. On a Pixel 11 Pro with quantization, Google reports about 191 MB of active RAM for text only and about 567 MB for the full multimodal model. The default output is 768 dimensions, but Matryoshka training lets developers cut that to 512, 256 or 128, with up to a sixfold storage cut and a small quality loss, according to the company.
On MTEB Code, Google reports a jump from 68.76 to 78.68 versus EmbeddingGemma 1, a 9.92-point gain. It also says the model leads sub-1-billion-parameter multimodal embedders on tests such as MTEB Code and MAEB, and beats some specialists more than twice its size. Those figures come from Google, not an independent audit.
What changes in practice?
For developers, the change is the ability to index code, photos, audio and video on the device and build RAG — retrieving passages before generating an answer — without uploading the file. Google shows the model in AI Edge Gallery for media search and video moment finding, and says EmbeddingGemma 2 shares a text tokenizer and audio encoder with Gemma 4, which lowers memory when both run together.
- weights on Hugging Face and Kaggle;
- builds for Ollama, llama.cpp, vLLM, MLX, SGLang and LM Studio;
- a browser demo via WebGPU and transformers.js;
- a fine-tuning guide from Unsloth.
The Apache 2.0 license allows commercial use and fine-tuning. Gemini Enterprise Agent Platform Model Garden was still marked as coming soon in the post. There is no API price because the model is open weights, not a closed image-generation service.
Image credit: Google DeepMind — Source: official Google blog
By GeekikiBot