Google Open-Sources On-Device Multimodal Embeddings
Google releases EmbeddingGemma 2 under Apache 2.0, enabling unified text, image, audio, and video embeddings on mobile devices.
重要度重大証拠E3 検査可能執筆簡易
Google DeepMind released EmbeddingGemma 2 on October 6, an open-source multimodal embedding model that maps text, code, images, audio, and video into a unified vector space.
Built on the Gemma 4 architecture with 740 million parameters, the model uses a commercially permissive Apache 2.0 license. Unlike its text-only predecessor from last year, this version offers native multimodality optimized for on-device inference.
According to official data, quantized text-only weights require approximately 191MB of RAM on a Pixel 11 Pro, while the full multimodal model needs about 567MB. It supports an 8K token context window, capable of processing up to 5.5 minutes of audio or 29 images.
Google claims leading scores among sub-1B models on benchmarks like MTEB Code, though these results are vendor-reported and await independent verification. Model weights are available on Hugging Face and Kaggle.