EmbeddingGemma 2: An open, lightweight multimodal embedding model

blog.google

411 points by ilreb a day ago


simonw - a day ago

I really appreciate that EmbeddingGemma 2 is under the Apache 2.0 license.

For embedding models in particular, I don't think it makes sense to use a closed, proprietary, hosted-only model.

Most applications of embedding models involve calculating thousands or even millions of embedding vectors and storing them for later comparison.

If your model is proprietary, the vendor is likely someday going to decide to stop offering that model. They'll have a better model to replace it, but you still need to pay to re-calculate those millions of stored existing vectors.

(In April 2024 OpenAI offered to "cover the financial cost of users re-embedding content with these new models" - https://openai.com/index/gpt-4-api-general-availability/ - but I don't think that's something we can rely on from every provider.)

Notably, I don't want to host the model myself. I'd much rather pay a provider for a hosted model while knowing that if they ever stop hosting it I can run the open weights version myself - or find another vendor who can do that for me.

Nautman - a day ago

It's also very neat that this can be used for "Jev"-like tasks with text and image.

https://developers.google.com/edge/mediapipe/solutions/decis...

minimaxir - a day ago

Finally. I was getting annoyed that there's been an inflection point in how LLMs/agents work but there hasn't been a good moderate-size embeddings model, and this one is multimodal too! 270M for text only is great compared to older embedding models, and a total 440M for text + vision is also fair.

I also may or may not have a tool for much faster local embedding creation that I calibrated for EmbeddingGemma but didn't want to release until a better embedding model came along.

flockonus - a day ago

Hats off to google for offering OSS (or at least open weights + license) a model that would be probably pretty closed to what they would ship in their Android phones.

jraedisch - 5 hours ago

First (Claude based) testing confirms not better or faster if one only needs text embeddings for retrieval or clustering. For music audio, MuQ-MuLan seems to remain the way to go? (Multimodal is ofc very useful for many other purposes.)

antonyragleap - 12 hours ago

Lightweight + Apache 2.0 is a great combo. Most multimodal embeddings are too heavy for on-device / self-hosted use. Nice to see Google pushing this.

aabhay - a day ago

Note that unlike prior on device embedding models, this seems to be trained with MRL, not MatFormers, meaning you don’t get to shrink the model weights alongside the lower dimensional embeddings, unfortunately. Likely there’s not good research for how to do MatFormers for multimodal yet?

kaycebasques - 19 hours ago

The JetBrains post from a few days ago introduced to me the idea of using binary quantization rather than MRL. Would that work with EmbeddingGemma2 or is there some reason why the approach might be fundamentally incompatible? https://news.ycombinator.com/item?id=49956148

*summons a minimaxir*

dcl - a day ago

Would be good to see how it compares to the embedding models from https://www.voyageai.com/ for text. I have used these a few times in the past and have found them superior to the Qwen models compared to here.

ContinuityLab - 10 hours ago

Lightweight, privacy-first multimodal embeddings are a crucial building block for running reliable vector representations locally without relying on external APIs.

djoldman - a day ago

Parameter count split is interesting:

740M total (270M text, 170M vision, 300M audio)

Juvination - a day ago

So what are some use cases people have found for running these sized multimodals on their device? What is it accurate on, and what is the hallucination rate like?

anyg - 17 hours ago

Interesting that they have buried the trending decisions API as the third example and a passing reference in the summary.

This may gain more traction if they lead with multimodal input decision making.

nicman23 - 6 hours ago

would pluging that to immich or frigate make sense or too slow?

brokensegue - a day ago

why isn't this being compared to siglip2 (also from google)? because that one isn't fully multimodal? or because it's a different org/team?

sourcecodeplz - a day ago

for text, benchmarks are identical to the first EmbeddingGemma.

but you can use this new one and enable/disable what you don't need.

can keep only text for ex.

nowittyusername - a day ago

I'm considering adding this in my harness after some testing, this seems like a really nice embedding model!

teaearlgraycold - 20 hours ago

I think I caused this release because I just embedded a sizable corpus with the previous model two days ago! You’re welcome everyone.

sohamactive - a day ago

rag transformations would be legendary

ChristmasTomer - 3 hours ago

[flagged]

favori995749721 - 17 hours ago

[flagged]

chriskr7 - 18 hours ago

[flagged]