EmbeddingGemma 2's vision tower mapped images and text into a 768-dim space; in-browser search using WebGPU with ruNNtime in TypeScript for GPU-based image-text retrieval
Read the original at www.reddit.com→EmbeddingGemma 2 came out this week. It maps images and text into one 768-dim space, so you can search photos by describing them. I ported its text and vision towers to ruNNtime, a WebGPU inference library in...
Original headline: "Image-text retrieval with EmbeddingGemma 2's vision tower, running in the browser on WebGPU"
Coverage timeline
- Oct 7, 09:26 UTC r/LocalLLaMA lead source Image-text retrieval with EmbeddingGemma 2's vision tower, running in the browser on WebGPU