Next upPhysical AI VC <> Founders Pitch Night #SFTechWeek @Mission Robotics
News

Google releases EmbeddingGemma 2 for on-device multimodal search

Google's 740-million-parameter EmbeddingGemma 2 maps text, code, images, audio and video into one embedding space for local search and retrieval.

D
Oct 6, 2026 · 2 min read

Google released EmbeddingGemma 2 on Oct. 6, an open 740-million-parameter model that maps text, code, images, audio and video into a shared embedding space.

The model gives developers one system for multimodal search, retrieval-augmented generation, classification and clustering on consumer devices. Generating embeddings locally can keep source material on the device and allow retrieval without a network connection. That expands Google’s local embedding work from text to several media types, although the supplied evidence does not establish performance across non-Google hardware or large real-world indexes.

Embeddings are numerical representations that place related content close together, allowing a search system to match a query with relevant material even when the words or media format differ. EmbeddingGemma 2 produces 768-dimensional vectors across its supported media types. Google’s model card lists a 270-million-parameter text component, a 170-million-parameter vision encoder and a 300-million-parameter audio encoder. Developers can omit encoders they do not need.

Google said quantized weights used about 191 MB of active memory for text-only operation and about 567 MB for the full model on a Pixel 11 Pro. Those figures are Google’s measurements, not independently reproduced results. The company also said Matryoshka Representation Learning allows the 768-dimensional vectors to be shortened to 512, 256 or 128 dimensions, reducing vector storage by as much as sixfold.

In Google’s full-precision evaluation, EmbeddingGemma 2 scored 78.68 on MTEB Code, compared with 68.76 for EmbeddingGemma 1, a 9.92-point increase. The model card also lists results across image, visual-document, video, speech and audio embedding suites. Google claims the model leads multimodal embedding models below one billion parameters on selected tests and can match or outperform some larger specialist models, but the research used for this article found no independent replication.

The model has an 8,192-token context window. Google says that capacity can hold about 5.5 minutes of audio, 29 images at its default image budget or 58 video frames at the default sampling rate when a single modality uses the available context. EmbeddingGemma 2’s local retrieval design is distinct from Google DeepMind’s planned server-side memory for Private AI Compute.

Google released the weights under the Apache 2.0 license through Hugging Face and Kaggle. Google’s developer guide documents use with sentence-transformers and other tools.

More news