// AI NEWS · MODEL RELEASE

Google Releases EmbeddingGemma 2, a 740M Open Multimodal Embedding Model Built to Run On-Device

Google DeepMind shipped EmbeddingGemma 2 on October 6, 2026 — an Apache-2.0 model that maps text, code, images, video, and audio into one embedding space and runs inside a few hundred megabytes of RAM on a phone.

2026-10-06 · by Moe Ameen

What happened

Google DeepMind released EmbeddingGemma 2 on October 6, 2026 — the successor to the original EmbeddingGemma, this time rebuilt on the Gemma 4 architecture and extended from text-only to multimodal. Where the first version embedded text, EmbeddingGemma 2 maps text, code, images, video, and audio into a single shared 768-dimensional embedding space, so a search across all of those modalities can run against one index. It ships under a commercially permissive Apache 2.0 license with weights on Hugging Face and Kaggle.

The design is modular and small by intent. The full multimodal model is about 740 million parameters, built from a 270M text-and-code core plus optional vision (170M) and audio (300M) encoders you add only if you need them. The context window grew to 8,192 tokens — four times the original — and the model carries over Matryoshka Representation Learning, so you can truncate the native 768-dimension vector down to 512, 256, or 128 dimensions to trade a little quality for much smaller storage. Google reports that 256-dimensional vectors keep roughly 95% of full retrieval quality on image, video, and speech.

The headline is that it runs on consumer hardware. Google cites text-only inference using about 191MB of active RAM on a Pixel 11 Pro, and the full multimodal suite around 567MB, with up to a 6x storage reduction for on-device vector databases via the smaller embedding dimensions. On quality, EmbeddingGemma 2 posts a 9.92-point jump on the MTEB Code benchmark (from 68.76 to 78.68) over its predecessor and is positioned as best-in-class among sub-1B multimodal embedders on benchmarks including MTEB Code and the Massive Audio Embedding Benchmark. One important clarification for creators: this is an embedding model, not a generative one — it turns content into vectors for search, routing, and retrieval; it writes no copy and renders no media.

Why it matters for creators

  • On-device retrieval gets real. A private semantic search over your own transcripts, scripts, captions, and footage can now run locally — on a phone or laptop — with no per-query API bill and nothing leaving the device.
  • One index for every modality. Because text, images, video, and audio share an embedding space, you can search a mixed content archive ("find the clip where I talked about pricing") instead of maintaining a separate index per format.
  • It finds; it does not make. EmbeddingGemma 2 is pure infrastructure — it surfaces the right source material, but it produces no post, caption, clip, or schedule. The creative and publishing work still has to happen somewhere.
  • The Matryoshka dimensions matter for scale. Truncating to 256 or 128 dims shrinks a vector database several-fold with minimal quality loss, which is what makes indexing a large back catalog on modest hardware practical.
  • Apache 2.0 means you can build on it commercially. Agencies and SaaS teams can embed it into their own RAG and content-intelligence pipelines without a license fee, paying only for the hardware that runs it.

How to act on this with Kompozy

The most useful way to read this release is as a new front door to your own content library. EmbeddingGemma 2 lets you index every transcript, caption, past post, and video frame you own and then ask it, in plain language, "what have I already said about X, and where is my best clip on it?" That is a genuine unlock for content intelligence — but the model stops at the answer. It hands you a ranked list of source material and nothing else: no script, no carousel, no edited short, no scheduled feed.

That is exactly where [Kompozy](/) picks up. Use EmbeddingGemma 2 (or a RAG tool built on it) to surface the strongest angle and the exact source to reuse, then drop that source into Kompozy and let it generate the finished content the model can't — a [Persona Short](/glossary/persona-shorts) with a face-locked avatar, [Clipped Shorts](/) from a longer talk, a brand-exact [Carousel via HyperFrames](/glossary/hyperframes), quote graphics, a blog, and a newsletter — all held to one voice by the [Persona Brief](/glossary/persona-brief). Then [Autopilot](/glossary/autopilot) reframes each asset per platform and publishes across eight social platforms plus blog and email behind a per-post review. The split is clean: EmbeddingGemma 2 is the retrieval brain that tells you what to make; Kompozy is the engine that makes and ships it everywhere.

Quick takeaways

  • Google DeepMind released EmbeddingGemma 2 on October 6, 2026 under an Apache 2.0 license, with weights on Hugging Face and Kaggle.
  • It is a ~740M-parameter multimodal embedding model (270M text/code core + optional 170M vision and 300M audio encoders) built on Gemma 4, mapping text, code, images, video, and audio into one 768-dim space.
  • It runs on-device — about 191MB RAM text-only and ~567MB full multimodal on a Pixel 11 Pro — with an 8,192-token context window and Matryoshka truncation to 512/256/128 dims.
  • It is for semantic search, routing, and retrieval — not generation. It embeds content; it writes and renders nothing.
  • Pair it with Kompozy: let the model find your best source material, then let Kompozy generate and publish the finished posts across nine platforms.

Frequently asked questions

What is EmbeddingGemma 2?

EmbeddingGemma 2 is an open (Apache 2.0) embedding model from Google DeepMind, released October 6, 2026. It is about 740M parameters, built on Gemma 4, and maps text, code, images, video, and audio into one shared 768-dimensional embedding space for semantic search, retrieval, and RAG. It is designed to run on-device.

How is EmbeddingGemma 2 different from the first EmbeddingGemma?

The original was a smaller, text-only embedding model. EmbeddingGemma 2 is multimodal — it embeds code, images, video, and audio as well as text — is built on the Gemma 4 architecture, has a 8,192-token context window (four times the original), and posts a large gain on code retrieval (MTEB Code rose from 68.76 to 78.68).

Can EmbeddingGemma 2 write posts or make videos?

No. It is an embedding model, not a generative one. It converts content into vectors so you can search and retrieve it; it does not write copy, generate images, edit video, or publish. To turn the source material it surfaces into finished posts you pair it with a content engine like Kompozy.

Why does on-device embedding matter?

Running the model locally means your content index — transcripts, captions, footage — never leaves your device, there is no per-query API cost, and search works offline. With Matryoshka truncation to 256 or 128 dimensions, the vector database also shrinks several-fold, so indexing a large archive on a phone or laptop becomes practical.

Related news

← All AI news · Get started →