Google releases EmbeddingGemma 2, an open embedding model that maps text, images, audio and video into one vector space
Announced October 6, 2026
What happened
Google DeepMind has released EmbeddingGemma 2, a lightweight open multimodal embedding model. According to the Hugging Face Transformers v5.19.0 release notes, it is based on the Gemma 4 architecture and encodes text, images, audio and video, alone or combined, into a shared 768-dimensional vector space. It uses Matryoshka Representation Learning, so embeddings can be truncated to 512, 256 or 128 dimensions. Unused vision and audio towers can be disabled at load time to save memory.
Why it matters
Transformers and Ollama (MLX) supported it on release day, so it can be used right away for cross-modal search and RAG. Simon Willison noted that its Apache 2.0 license reduces the risk of having to recompute embeddings.
Sources
- Official EmbeddingGemma 2: an open, lightweight multimodal embedding model Google DeepMind
- Official Release v5.19.0 Hugging Face
- Community EmbeddingGemma 2 Simon Willison
More about Google
-
Google releases image generation and editing model Nano Banana 2.1, cutting generation cost by about half (report)
According to AI Times, Google released the image generation and editing model 'Nano Banana 2.1' on the 6th (local time). The report says...
-
Anthropic launches public beta of Claude for Google Workspace, usable directly in Docs, Sheets and Slides (report)
According to AI Times, Anthropic released 'Claude for Google Workspace' as a public beta on all paid Claude plans on the 6th (local...
-
Google signs 3.59 GW nuclear power deal with Constellation to power AI data centers (report)
According to AI Times, Google signed a 3.59 GW nuclear power supply agreement with Constellation Energy on the 6th (local time),...