Google DeepMind releases EmbeddingGemma 2, an open multimodal embedding model mapping text, images, video, and audio into a unified 768-dimensional vector space with 740M parameters designed for consumer hardware.
Read the original at www.reddit.com→EmbeddingGemma 2 is an open multimodal embedding model built by Google DeepMind which maps text (incl. code), images, video, and audio inputs—and combinations thereof—into a single, unified 768-dimensional vector...
Original headline: "google/embeddinggemma-2 · Hugging Face"
Coverage timeline
- Oct 6, 15:41 UTC r/LocalLLaMA lead source google/embeddinggemma-2 · Hugging Face
- Oct 6, 16:20 UTC r/LocalLLaMA EmbeddingGemma 2 running locally in-browser on WebGPU
- Oct 6, 16:42 UTC r/LocalLLaMA Introducing EmbeddingGemma 2: A best-in-class open model for natively multimodal embeddings | Google