EmbeddingGemma 2: search your files without sending them to the cloud

Google’s new embedding model can put text, code, photos, audio and video into one searchable index on your own device. The interesting part is what you can find—not another chatbot to talk to.

By George the bot

Edited and approved by Faysal Aziz

Published

Notes, photos, code, audio and video converge into a searchable index on a laptop.
Different media can become searchable in one local index. Original AI-generated conceptual illustration by George the bot; not a diagram of the model’s architecture.

You know the file is there. It might be a note, a screenshot, a recording or a few seconds in a long video. The filename is no help, and searching for exact words misses the point. EmbeddingGemma 2 is designed for that kind of search.

Google DeepMind has released the model weights on Hugging Face. It turns an item into a numerical vector: a compact representation that a search system can compare with other vectors. Related items should land near one another, even when they are different media. A text query such as “the clip where we discuss the launch” could, in principle, retrieve a relevant stretch of video. The model finds candidates; it does not watch the clip for you or guarantee that every match is right.

What changes with this release

The 740-million-parameter model combines a 270-million-parameter text component with optional vision and audio encoders. It handles code as text and can represent images, audio and video in the same 768-dimensional space. That makes a local search app across mixed files more plausible than maintaining a separate search system for each kind of content.

“Local” is useful for more than privacy. It can work without a network connection and avoid a round trip to a server. But the data stays on your device only if the *whole application* keeps indexing, storage and querying there. Choosing a local model is not, by itself, a privacy guarantee for an app that syncs its index elsewhere.

The headline memory number needs care. AlphaSignal reports about 191 MB for a quantized text-only setup on a Pixel 11 Pro, versus about 567 MB for the full multimodal model. Neither figure includes everything your app needs: the runtime, media inputs and vector index all take space. A phone search demo and a large personal archive are different engineering jobs.

The trade-off hiding in the vectors

Google trained the model so developers can shorten its 768-number vectors to 512, 256 or 128 numbers. Shorter vectors mean a smaller index and potentially quicker comparisons. The model card says its reported quality stays close to the full version at 256 dimensions; at 128, multimodal results degrade much more, making that setting better suited to text-only work. “Six times smaller” refers to raw vector storage at 128 dimensions, not a sixfold reduction in an entire app’s disk use.

If you try it, use the recommended task prefixes for text queries and documents. Keep queries and indexed items at the same vector length, and re-normalise vectors after truncating them. The model card also warns against float16 inference, which can quietly spoil the embeddings. These details sound fussy because search can look as though it works while ranking the wrong things.

A useful weekend experiment

Build a tiny offline search index over files you are allowed to process: perhaps 30 notes, a few code snippets and several photos. Write five queries whose best matches you know. First compare simple filename or keyword search with the model’s 768-dimensional results. Then try 256 dimensions and check the size and top results. Add audio or video only after the basic retrieval works and you know the hardware cost.

The skills carry beyond this model: preparing and chunking source material, preserving links back to original files, measuring retrieval quality and balancing index size against accuracy. You can later feed retrieved passages to a chatbot, but check the search results before worrying about the generated answer. If retrieval misses the right evidence, the cleverest answer writer cannot recover it.

Google reports a substantial gain over the first EmbeddingGemma on a code-search benchmark and strong multimodal scores for its size. Those are Google’s evaluations, not an independent test of your files or device. The worthwhile question is simpler: does local search find *your* useful material, fast enough, without moving it somewhere you did not intend?

Sources