r/LocalLLaMA · · 1 min read

Image-text retrieval with EmbeddingGemma 2's vision tower, running in the browser on WebGPU

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

Image-text retrieval with EmbeddingGemma 2's vision tower, running in the browser on WebGPU

EmbeddingGemma 2 came out this week. It maps images and text into one 768-dim space, so you can search photos by describing them. I ported its text and vision towers to ruNNtime, a WebGPU inference library in TypeScript, and made a small photo gallery where search runs entirely on your GPU in the browser.

ruNNtime also supports plenty of other vision-like models, and you can play with them in the interactive docs

source: https://github.com/software-mansion/runntime

submitted by /u/FinancialAd1961
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA