Load EmbeddingGemma 2 in plain float16 and you can get NaNs – or ruined vectors with zero error. Activations overflow the format. That trap already burned early testers; the model card is blunt: do not use float16.
If you’ve stitched a text embedder, a CLIP-style vision model, and a separate audio tower just to search one local folder of notes, screenshots, and voice memos, you know the tax. Three downloads. Misaligned spaces. Late fusion glue. Your laptop fans spin up and the index still feels fragile.
EmbeddingGemma 2 shipped October 6, 2026 from Google DeepMind – open weights, Apache 2.0, on Hugging Face and Kaggle. 740M total, modular, built on the Gemma 4 line. Text, code, images, video, audio (or mixes) land in one 768d space with Matryoshka heads down to 512/256/128. The practical gain: offline multimodal RAG without the Frankenstein stack. Day-one chatter chased the code MTEB jump (+9.92 vs v1) and on-device privacy; the part that actually changes a build is selective loading plus the failure modes most launch posts bury.
Funny how often “just embed everything” turns into a weekend of alignment scripts. Same folder, same human intent – different vector universes. That mismatch is what this model is trying to erase.
Why separate local embedders still hurt
Multiple indexes. Approximate cross-modal scores. Memory that adds up on consumer hardware. Closed multimodal APIs that ship private photos or meeting audio off-device. The original EmbeddingGemma stayed text-only with a 2K context and still passed 20M downloads – useful, not enough for a media-heavy project folder.
This release stays small while adding native vision and audio encoders that project into the same space. A text query can rank a video clip or voice note with plain cosine. Shared context is 8192 tokens as of the October 2026 launch specs – tight once you interleave high-res images with long text, not an infinite bucket.
Selective load + task prompts
Install the stack that actually supports it:
pip install -U "sentence-transformers>=6.1.0" transformers pillow soundfile torchcodec
dtype first. bfloat16 when the GPU supports it, else float32 – never float16.
import torch
from sentence_transformers import SentenceTransformer
dtype = torch.bfloat16 if torch.cuda.is_bf16_supported() else torch.float32
model = SentenceTransformer(
"google/embeddinggemma-2",
config_kwargs={"vision_config": None, "audio_config": None}, # text/code 270M
model_kwargs={"torch_dtype": dtype}
)
Turns out you can peel the model like an onion: text/code 270M, add vision for 440M, other mixes up to 570M, full 740M. Empty config_kwargs loads everything. All variants share the space, so a text-only corpus does not need re-embedding when you later flip vision on.
Text retrieval wants prompts. Skip them and quality drops for free:
query_emb = model.encode("find the auth middleware", prompt_name="CodeRetrieval")
doc_emb = model.encode("title: auth.py | text: def require_auth...", prompt_name="Document")
print(model.similarity(query_emb, doc_emb))
Prompt names that matter in practice: SearchQuery, Document, CodeRetrieval, QuestionAnswering, Classification, Clustering, SentenceSimilarity. Media takes dicts – no prefix.
Local mixed media + code index
Project folder: README, source, demo.mp4, design PNG. One encode call on an interleaved card:
emb = model.encode({
"text": "New login flow. Screen capture of the form. Demo: ",
"image": ["login.png"],
"video": "demo.mp4"
})
Query: “how does the login form validate email”. That joint vector ranks beside pure code chunks. Offline on a laptop. Google’s quantized figures (as of launch): ~191 MB active RAM text-only, ~567 MB full on a Pixel-class device. Put vectors in something local (Qdrant works) and keep generation on a small Gemma 4 if you want the full private loop.
Video defaults to 1 fps – override with processing_kwargs when you need denser frames. Audio path expects 16 kHz mono; resampling is automatic.
Pro tips that change ranking
After any truncate_dim, L2-normalize again. Slicing a unit vector breaks unit length. Cosine values still look plausible; order collapses. Use normalize_embeddings=True with truncate_dim=256 (or 128/512) and keep query/corpus dims identical.
256d holds most text/code quality and the bulk of multimodal retrieval while cutting storage 3× (HF truncation table). 128d is fine for cheap text shortlists. Multimodal? MMEB overall falls 59.01 → 45.65 at 128d per the model card – measure on your own images and audio before you standardize there.
Is 128d ever worth it for a photo+voice corpus, or do you just pay the 256d storage and sleep better? Depends how brutal your recall curve looks past the first ten hits.
The catch is the shared 8192-token budget. Default image cost sits near ~280 tokens, a video frame ~140, audio about 25 tokens per second. Interleave three hefty images with a long README and the per-modality headroom shrinks fast. Also: do not mix EmbeddingGemma 1 vectors into a v2 index. Different spaces.
FAQ
Can I run EmbeddingGemma 2 fully offline on a phone or laptop?
Yes. Open weights, Apache 2.0. Quantized text-only sits under ~200 MB active RAM per Google’s launch numbers. Full multimodal is realistic on recent consumer hardware.
Do I need different models for text-to-image vs image-to-audio search?
No. Everything maps into the same 768d space. Text query, image gallery, audio clip – plain cosine after normalization. Interleaved inputs already return one joint vector, so you are not chaining specialist models or calibrating fusion weights.
What’s the biggest quality drop I should watch for?
People chase the 128d storage win and miss the quieter killers. float16 can emit silent garbage – bfloat16 or float32 only. Task prefixes on the text side move scores more than folks expect. Treat 128d multimodal as guilty until your own recall numbers say otherwise; 256d is the safer default. (Details and the MMEB drop sit in the pro-tips block above.)
Clone the Hugging Face repo, load the text-only config, embed ten of your own files with the right prompt_name, and check a few cross-modal scores. That five-minute pass beats another benchmark screenshot. Specs and multimodal guides: Gemma docs. Launch context: Google blog.