01 What happened

On October 6, 2026, Google released EmbeddingGemma 2, a 740-million-parameter model built on Gemma 4 that maps text, code, images, audio, and video into a shared embedding space for on-device retrieval. The weights are now on Hugging Face and Kaggle; Gemini Enterprise Agent Platform Model Garden availability is announced as coming soon.

02 Key details

  • EmbeddingGemma 2 has 740M parameters and handles text, code, images, video, and audio. Text-only workloads need as little as 270M parameters, with optional vision (170M) and audio (300M) encoders for full multimodal support.
  • With Matryoshka Representation Learning, developers can truncate output vectors from 768 down to 512, 256, or 128 dimensions, providing up to 6x storage and memory reduction in local vector databases.
  • Quantized on a Google Pixel 11 Pro, Google reports as little as ~191MB of active RAM for text-only weights and ~567MB for the full multimodal model.
  • Google reports a 9.92-point gain in MTEB Code, from 68.76 to 78.68, and an 8K token context window 4x larger than EmbeddingGemma 1, supporting up to 5.5 minutes of audio, 29 images, 58 video frames, or interleaved combinations.

03 Why it matters

The release puts multimodal embeddings to work on local hardware. A modular design lets text-only workloads run with as little as 270M parameters, and Matryoshka Representation Learning cuts vectors from 768 to 512, 256, or 128 dimensions, reducing storage and memory by up to 6x. Google also reports an 8K token context window and a 9.92-point gain in MTEB Code, from 68.76 to 78.68; these are company-reported figures, not independently verified.

04 Who it matters to

Local search developers, RAG system developers and AI agent developers.

Original sourceGoogle