Google DeepMind shipped EmbeddingGemma 2 on October 6, 2026 — an Apache-2.0 model that maps text, code, images, video, and audio into one embedding space and runs inside a few hundred megabytes of RAM on a phone.
2026-10-06 · by Moe Ameen
Google DeepMind released EmbeddingGemma 2 on October 6, 2026 — the successor to the original EmbeddingGemma, this time rebuilt on the Gemma 4 architecture and extended from text-only to multimodal. Where the first version embedded text, EmbeddingGemma 2 maps text, code, images, video, and audio into a single shared 768-dimensional embedding space, so a search across all of those modalities can run against one index. It ships under a commercially permissive Apache 2.0 license with weights on Hugging Face and Kaggle.
The design is modular and small by intent. The full multimodal model is about 740 million parameters, built from a 270M text-and-code core plus optional vision (170M) and audio (300M) encoders you add only if you need them. The context window grew to 8,192 tokens — four times the original — and the model carries over Matryoshka Representation Learning, so you can truncate the native 768-dimension vector down to 512, 256, or 128 dimensions to trade a little quality for much smaller storage. Google reports that 256-dimensional vectors keep roughly 95% of full retrieval quality on image, video, and speech.
The headline is that it runs on consumer hardware. Google cites text-only inference using about 191MB of active RAM on a Pixel 11 Pro, and the full multimodal suite around 567MB, with up to a 6x storage reduction for on-device vector databases via the smaller embedding dimensions. On quality, EmbeddingGemma 2 posts a 9.92-point jump on the MTEB Code benchmark (from 68.76 to 78.68) over its predecessor and is positioned as best-in-class among sub-1B multimodal embedders on benchmarks including MTEB Code and the Massive Audio Embedding Benchmark. One important clarification for creators: this is an embedding model, not a generative one — it turns content into vectors for search, routing, and retrieval; it writes no copy and renders no media.
The most useful way to read this release is as a new front door to your own content library. EmbeddingGemma 2 lets you index every transcript, caption, past post, and video frame you own and then ask it, in plain language, "what have I already said about X, and where is my best clip on it?" That is a genuine unlock for content intelligence — but the model stops at the answer. It hands you a ranked list of source material and nothing else: no script, no carousel, no edited short, no scheduled feed.
That is exactly where [Kompozy](/) picks up. Use EmbeddingGemma 2 (or a RAG tool built on it) to surface the strongest angle and the exact source to reuse, then drop that source into Kompozy and let it generate the finished content the model can't — a [Persona Short](/glossary/persona-shorts) with a face-locked avatar, [Clipped Shorts](/) from a longer talk, a brand-exact [Carousel via HyperFrames](/glossary/hyperframes), quote graphics, a blog, and a newsletter — all held to one voice by the [Persona Brief](/glossary/persona-brief). Then [Autopilot](/glossary/autopilot) reframes each asset per platform and publishes across eight social platforms plus blog and email behind a per-post review. The split is clean: EmbeddingGemma 2 is the retrieval brain that tells you what to make; Kompozy is the engine that makes and ships it everywhere.
EmbeddingGemma 2 is an open (Apache 2.0) embedding model from Google DeepMind, released October 6, 2026. It is about 740M parameters, built on Gemma 4, and maps text, code, images, video, and audio into one shared 768-dimensional embedding space for semantic search, retrieval, and RAG. It is designed to run on-device.
The original was a smaller, text-only embedding model. EmbeddingGemma 2 is multimodal — it embeds code, images, video, and audio as well as text — is built on the Gemma 4 architecture, has a 8,192-token context window (four times the original), and posts a large gain on code retrieval (MTEB Code rose from 68.76 to 78.68).
No. It is an embedding model, not a generative one. It converts content into vectors so you can search and retrieve it; it does not write copy, generate images, edit video, or publish. To turn the source material it surfaces into finished posts you pair it with a content engine like Kompozy.
Running the model locally means your content index — transcripts, captions, footage — never leaves your device, there is no per-query API cost, and search works offline. With Matryoshka truncation to 256 or 128 dimensions, the vector database also shrinks several-fold, so indexing a large archive on a phone or laptop becomes practical.