Google · 2026-10-06 · major
EmbeddingGemma 2 — Google's open 740M embedding model adds images, audio and video
EmbeddingGemma 2 is Google's new open embedding model. It puts text, code, images, audio and video in one shared vector space, has 740M parameters and an 8K context, and runs on phones. It is Apache 2.0.

One small open model that turns text, code, pictures, sound and video into vectors you can search together, on device.
Key specs
| Parameters | 740M |
|---|---|
| Mteb code v1 | 78.68 |
Quick facts
| Maker | |
|---|---|
| Size | 740M (270M text + 170M vision + 300M audio) |
| Context window | 8K tokens (4x EmbeddingGemma 1) |
| Embedding size | 768, truncatable to 512 / 256 / 128 |
| Inputs | Text, code, images, audio, video |
| License | Apache 2.0 |
| Where | Hugging Face, Kaggle |
What is it?
EmbeddingGemma 2 is Google's second open embedding model, released on October 6, 2026. The first version handled only text; this one also takes images, audio and video and maps them into the same 768-dimension space. It has 740 million parameters, an 8K-token context window and an Apache 2.0 license.
How does it work?
The design is modular: a 270M text backbone built on the Gemma 4 architecture, plus an optional 170M vision encoder and 300M audio encoder. In one pass the model can read up to 5.5 minutes of audio, 29 images or 58 video frames. Matryoshka training means the output vector can be cut to 512, 256 or 128 dimensions for up to 6x less storage.
Why does it matter?
Most on-device search and RAG stacks need a separate model per media type. With one multimodal embedder, an app can find a video clip from a voice memo or search hours of audio locally. Code retrieval also improves: the MTEB Code score rises from 68.76 to 78.68, a 9.92-point gain over the first EmbeddingGemma.
Who is it for?
developers building search, RAG and on-device apps
Frequently asked questions
- Can EmbeddingGemma 2 run on a phone?
- Yes. Google measured EmbeddingGemma 2 on a Google Pixel 11 Pro with quantization at about 191MB of active RAM for text only and about 567MB with the full multimodal encoders. On-device deployment goes through MediaPipe and LiteRT, and the model is also in the Google AI Edge Gallery.
- Which frameworks support EmbeddingGemma 2?
- Google lists transformers, sentence-transformers, MLX, vLLM, llama.cpp, SGLang, Ollama and LM Studio as supported runtimes for EmbeddingGemma 2. Weights are on Hugging Face and Kaggle, and the model card shows a sentence-transformers example with separate SearchQuery and Document prompts.
- How does EmbeddingGemma 2 compare to the first EmbeddingGemma?
- EmbeddingGemma 2 adds image, audio and video input to the text-only original, grows to 740M parameters and has a context window four times larger at 8K tokens. On code retrieval the MTEB Code score goes from 68.76 to 78.68. Google says it beats some specialist models more than twice its size but names none.
- Is EmbeddingGemma 2 free for commercial use?
- EmbeddingGemma 2 is released under the Apache 2.0 license, which allows commercial use. The model card adds that use is governed by the Gemma Prohibited Use Policy and that developers should build their own safeguards at the application level.
Try it
SentenceTransformer("google/embeddinggemma-2")