HeadlinesBriefing HeadlinesBriefing.com

Embedding Gemma 2: Multimodal On-Device Model

Hacker News •
×

Embedding Gemma 2 is the most capable model for on-device multimodal embeddings, natively mapping text, images, audio, and video into a unified embedding space. Introduced by Google DeepMind, it expands on the original Embedding Gemma, which saw over 20 million downloads. Built on the Gemma 4 architecture under an Apache 2.0 license, it has 740 million parameters, optimized for on-device inference.

It achieves leading scores among sub-1B multimodal embedders on benchmarks like MTEB Code and MAEB, matching or outperforming larger models. Modular design requires as little as 270M parameters for text-only tasks, with optional vision and audio encoders. Using Matryoshka Representation Learning, output vectors can be truncated from 768 down to 128 dimensions, reducing storage by up to 6x.

Optimized for edge hardware, it runs with as little as ~191MB active RAM on a Google Pixel 11 Pro for text-only, and ~567MB for full multimodal. It features an 8K token context window, 4x larger than before, processing up to 5.5 minutes of audio, 29 images, or 58 video frames. It shows a 9.92-point improvement on code performance, from 68.76 to 78.68 in MTEB Code.

Embedding Gemma 2 enables fully on-device semantic search, routing, and retrieval, ensuring data privacy and low latency. It pairs with generative models like Gemma 4 for on-device RAG pipelines, sharing tokenizer and audio encoder for a unified pipeline with lower memory footprint.

Source: Hacker News · Summarized by HeadlinesBriefing