On 2026-10-06 Google DeepMind released EmbeddingGemma 2, a 740M-parameter Apache-2.0 embedding model built on Gemma 4 that maps text, code, images, audio and video into one 768-dim space with an 8K context. MTEB Code rises from 68.76 to 78.68, and quantized text-only inference needs ~191MB RAM on a Pixel 11 Pro. Hugging Face Transformers v5.19.0 supports it natively.
Key Takeaways
- ✓Size: 740M params, modular 270M text backbone + 170M vision + 300M audio encoders, loadable on demand
- ✓Code retrieval: MTEB Code 68.76 → 78.68 (+9.92); Google claims leading sub-1B multimodal scores on MTEB Code and MAEB
- ✓Storage: Matryoshka embeddings truncate 768 → 512/256/128 dims, up to 6x smaller local vector stores
- ✓On-device: ~191MB RAM text-only / ~567MB full multimodal (quantized, Pixel 11 Pro); 8K context (4x v1) covers 5.5 min audio, 29 images or 58 video frames
- ✓Ecosystem: Apache 2.0 weights on Hugging Face and Kaggle; supported in Transformers v5.19.0, plus sentence-transformers, vLLM, llama.cpp, SGLang, Ollama and MLX per Google

Key Decision Metrics at a Glance
Turn your technical choice into a development budget
Compare 40 dev plans & simulate token costs vs $20/mo subscriptions
Project Links & Resources
Direct AccessIn-Depth Technical Analysis
Google DeepMind released EmbeddingGemma 2 on 2026-10-06, the successor to the text-only EmbeddingGemma (20M+ downloads). Built on the Gemma 4 architecture with 740M parameters under Apache 2.0, it embeds text, code, images, audio and video, alone or interleaved, into a shared 768-dim space. The design is modular: 270M for text-only use, with optional 170M vision and 300M audio encoders that can be disabled at load time. Matryoshka Representation Learning lets you truncate vectors to 512, 256 or 128 dims (up to 6x storage savings), and the 8K context (4x v1) fits about 5.5 minutes of audio, 29 images or 58 video frames. Google reports MTEB Code rising from 68.76 to 78.68 and leading sub-1B multimodal results on MTEB Code and MAEB; quantized, it needs about 191MB RAM text-only and 567MB fully multimodal on a Pixel 11 Pro. Weights are on Hugging Face and Kaggle, Transformers v5.19.0 adds native support, and Google lists sentence-transformers, vLLM, llama.cpp, SGLang, Ollama, MLX, LM Studio and transformers.js as ready runtimes. No independent third-party benchmarks yet.
Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.
Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.