I really appreciate that EmbeddingGemma 2 is under the Apache 2.0 license.
For embedding models in particular, I don't think it makes sense to use a closed, proprietary, hosted-only model.
Most applications of embedding models involve calculating thousands or even millions of embedding vectors and storing them for later comparison.
If your model is proprietary, the vendor is likely someday going to decide to stop offering that model. They'll have a better model to replace it, but you still need to pay to re-calculate those millions of stored existing vectors.
(In April 2024 OpenAI offered to "cover the financial cost of users re-embedding content with these new models" - https://openai.com/index/gpt-4-api-general-availability/ - but I don't think that's something we can rely on from every provider.)
Notably, I don't want to host the model myself. I'd much rather pay a provider for a hosted model while knowing that if they ever stop hosting it I can run the open weights version myself - or find another vendor who can do that for me.
IMO embedding models are some of the most likely models you’d want to self host because they are much more useful out of the box for batch or stream workloads than LLMs. You can just rent a GPU for $1-2/hr (you don’t need a B300) and run across your entire dataset.
If you are only interested in recommendations, clustering, etc and won’t have to do embedding at runtime, that’s it.
There is a lot of basic matrix stuff you can do with embedding vectors that are much easier than their LLM equivalents. I can’t remember the name for the techniques but a major one is translating questions annout a subject eg “Who is Slim Shady?”->”Marshall Mathers”so that your input distribution at runtime matches your offline distribution. Much easier to do with open weight models.
Not really relevant in the case of embedding models. We might note that compressing a lossless PNG to 99% quality JPEG will also give you a different embedding. You don't care about bit for bit equality, just the distance in embedding space. And the same text embedded on CPU vs GPU will give very very similar embeddings.
I got this running locally, and funnily enough the exact example they have for classifying "Cancel my flight and refund my credit card immediately." failed. It said the request does not involve a payment, charge, or refund, with a p(true) of 0.22 (where true means it is financial). Laya was significantly more accurate (in this case) and almost as fast.
Dang that is a neat trick! Embed the answers then dot product with the question embedding. All kinds of fun linear algebra tricks (centering, rescaling etc.) to improve outcomes.
Finally. I was getting annoyed that there's been an inflection point in how LLMs/agents work but there hasn't been a good moderate-size embeddings model, and this one is multimodal too! 270M for text only is great compared to older embedding models, and a total 440M for text + vision is also fair.
I also may or may not have a tool for much faster local embedding creation that I calibrated for EmbeddingGemma but didn't want to release until a better embedding model came along.
Not just vision with video, but also audio, it really seems amazing.
I’m not sure how it can handle vicinity of pairs of embeddings with for example some words and the audio where they’re spoken or an image where the text is handled. Building local multimodal search with this would be amazing.
I’ve explored this stuff with CLIP and it’s interesting how image (but also audio) embedding carries both the clean “text” content information but also the stylistic and visual/audio tone information, the two can even kind of be linearly separated.
> Building local multimodal search with this would be amazing.
(lawyer here) — I’m curious: for what you’re describing, wouldn’t the machine need to be constantly running/updating the embeddings to take updated and new files into account? If so, how would that computation load compare to, say, Spotlight constantly updating its index?
The embedding model stays loaded in memory. It is used for turning your search keywords into embeddings.
The index you’re thinking of is made once per file.. then you compare and search in embeddings. Add/modify files = asynchronous updating or adding corresponding embeddings using the model in memory onto wherever you persist those embeddings(say SQLite)..
Also how you turn a file into one or more items is a separate question and would likely need tuning to circumstances, usually big documents are split (chunked) sometimes at paragraph or even more granularly. Where to optimally chunk alone is not easy. This also allows you to then search for a specific part of the document, at the expense of not taking the wide context into account, but embeddings usually struggle with too many tokens anyway.
Numbers update on my hypothetical tool after Opus 5.5 added support for image and audio input (plus video support via ffmpeg, why not) using my M3 Pro:
- Text: 78 embeddings/second on small texts
- Image: 4 embeddings/second
- Audio: 6 embeddings/second for 30 second chunks (which is how the encoder works)
- Video: 0.2 embeddings/second per minute of video (not entirely surprising since the encoder does 1fps).
Not bad numbers for a local laptop running a model this large, although the M3 Pro is a few years old and I suspect a M5 Ultra will 5x them at minimum.
Hats off to google for offering OSS (or at least open weights + license) a model that would be probably pretty closed to what they would ship in their Android phones.
Note that unlike prior on device embedding models, this seems to be trained with MRL, not MatFormers, meaning you don’t get to shrink the model weights alongside the lower dimensional embeddings, unfortunately. Likely there’s not good research for how to do MatFormers for multimodal yet?
The JetBrains post from a few days ago introduced to me the idea of using binary quantization rather than MRL. Would that work with EmbeddingGemma2 or is there some reason why the approach might be fundamentally incompatible? https://news.ycombinator.com/item?id=49956148
Lightweight, privacy-first multimodal embeddings are a crucial building block for running reliable vector representations locally without relying on external APIs.
Would be good to see how it compares to the embedding models from https://www.voyageai.com/ for text. I have used these a few times in the past and have found them superior to the Qwen models compared to here.
It's been awhile since I've been in the space, but Voyage was never a serious contender outside of super-niche business domains. I suspect that this compares favorably in 9X% of use cases
Vision is just processing a still image (why it's by far the smallest). Text requires dealing with the entropy of human language. Audio is meaningless without time.
So what are some use cases people have found for running these sized multimodals on their device? What is it accurate on, and what is the hallucination rate like?
I run offmetaedh.com off of them and this model is a drop in replacement for gemma embedding v1 and HQ-CLIP. Better quality, less RAM usage, faster embeddings, I'm fucking pumped for it.
The image search in the Edge Gallery app of theirs is really incredible. Blows their Screenshot app out of the water for what I wanted it for. I stopped using it when they quietly changed their wording around it suggesting it would use Gemini instead of Gemini Nano, which was the whole reason I was excited about it.
The main one is mapping images to text and visa versa, e.g. semantic search of images via text, where the images are encoded and the text question is encoded with the same model, then finding nearest neighbors.
I haven’t tried it yet, but, for law practice, I could imagine using it to search a case file for “undamaged roof before Hurricane Katrina” and “damaged roof after Hurricane Katrina” and being able to locate both deposition testimony and pertinent photographs in the body of evidence.
For embedding models in particular, I don't think it makes sense to use a closed, proprietary, hosted-only model.
Most applications of embedding models involve calculating thousands or even millions of embedding vectors and storing them for later comparison.
If your model is proprietary, the vendor is likely someday going to decide to stop offering that model. They'll have a better model to replace it, but you still need to pay to re-calculate those millions of stored existing vectors.
(In April 2024 OpenAI offered to "cover the financial cost of users re-embedding content with these new models" - https://openai.com/index/gpt-4-api-general-availability/ - but I don't think that's something we can rely on from every provider.)
Notably, I don't want to host the model myself. I'd much rather pay a provider for a hosted model while knowing that if they ever stop hosting it I can run the open weights version myself - or find another vendor who can do that for me.
If you are only interested in recommendations, clustering, etc and won’t have to do embedding at runtime, that’s it.
There is a lot of basic matrix stuff you can do with embedding vectors that are much easier than their LLM equivalents. I can’t remember the name for the techniques but a major one is translating questions annout a subject eg “Who is Slim Shady?”->”Marshall Mathers”so that your input distribution at runtime matches your offline distribution. Much easier to do with open weight models.
https://developers.google.com/edge/mediapipe/solutions/decis...
I also may or may not have a tool for much faster local embedding creation that I calibrated for EmbeddingGemma but didn't want to release until a better embedding model came along.
I’m not sure how it can handle vicinity of pairs of embeddings with for example some words and the audio where they’re spoken or an image where the text is handled. Building local multimodal search with this would be amazing. I’ve explored this stuff with CLIP and it’s interesting how image (but also audio) embedding carries both the clean “text” content information but also the stylistic and visual/audio tone information, the two can even kind of be linearly separated.
(lawyer here) — I’m curious: for what you’re describing, wouldn’t the machine need to be constantly running/updating the embeddings to take updated and new files into account? If so, how would that computation load compare to, say, Spotlight constantly updating its index?
The index you’re thinking of is made once per file.. then you compare and search in embeddings. Add/modify files = asynchronous updating or adding corresponding embeddings using the model in memory onto wherever you persist those embeddings(say SQLite)..
- Text: 78 embeddings/second on small texts
- Image: 4 embeddings/second
- Audio: 6 embeddings/second for 30 second chunks (which is how the encoder works)
- Video: 0.2 embeddings/second per minute of video (not entirely surprising since the encoder does 1fps).
Not bad numbers for a local laptop running a model this large, although the M3 Pro is a few years old and I suspect a M5 Ultra will 5x them at minimum.
*summons a minimaxir*
740M total (270M text, 170M vision, 300M audio)
Vision is just processing a still image (why it's by far the smallest). Text requires dealing with the entropy of human language. Audio is meaningless without time.
This may gain more traction if they lead with multimodal input decision making.
but you can use this new one and enable/disable what you don't need.
can keep only text for ex.