EmbeddingGemma 2: An open, lightweight multimodal embedding model
Points and comments are a snapshot, not live.
Google launches EmbeddingGemma 2, an open multimodal embedding model with 740M parameters.
EmbeddingGemma 2, built on Gemma 4, natively maps text, images, audio, and video into a unified embedding space. It has 740M parameters, with optional encoders for text-only (270M), vision (170M), and audio (300M) workloads. Released under Apache 2.0, it achieves top benchmarks among sub-1B models and supports Matryoshka Representation Learning for dynamic vector truncation. It runs on-device with as little as 191MB RAM for text-only and 567MB for full multimodal use, and features an 8K token context window. Available on Hugging Face and Kaggle.
What commenters are saying
Commenters widely welcome the release, praising the open Apache 2.0 license and multimodal capabilities. Several highlight the practical advantage of open weights for embedding models, noting that proprietary models risk vendor lock-in and re-embedding costs. One commenter reports real-world performance on an M3 Pro: 78 text embeddings/second, 4 images/second, 6 audio embeddings/second for 30-second chunks, and 0.2 video embeddings/second per minute. A lawyer sees potential for searching case files by text and images. Some question comparisons to Google's own SigLIP 2 and note the absence of MatFormers training for weight reduction alongside MRL.