Bring multimodal semantic search to the edge with EmbeddingGemma 2
EmbeddingGemma 2: Multimodal Semantic Search at the Edge
Overview
EmbeddingGemma 2 is a new 740M open-weight multimodal model that maps text, images, video, and audio into a unified vector space for privacy-first, on-device retrieval. It enables ultra-low-latency local solutions such as search-as-you-type media retrieval, keyframe video moments finding, and zero-shot intent routing. For developers building local search and media retrieval experiences, it reduces the latency and memory overhead of chaining separate image captioning, speech-to-text, and text-embedding models.
Key Features
- Unified Vector Space: Natively maps text, images, video frames, and audio into a single, unified vector space
- On-Device Decision Engine: Matches user inputs directly against classification labels and descriptions without any training data or fine-tuning, delivering instant, zero-shot intent routing in milliseconds
- Cross-Modal Representations: Supports text, vision, and audio modalities within a compact 740M parameter footprint
Performance & Hardware Requirements
| Specification | Detail |
|---|---|
| Model Size | 740M parameters |
| Text-Only Memory | ~191MB active RAM |
| Full Multimodal Memory | ~567MB on Google Pixel 11 Pro |
| Visual Embedding Latency | 37.3 ms (26.9 images per second) on MacBook M5 Pro GPU |
| Decision Engine Response Time | Less than 100ms |
Integration Options
- MediaPipe Tasks: Provides a turnkey solution for embedding and semantic retrieval across multiple platforms
- LiteRT: Optimizes performance across CPUs, GPUs, and NPU accelerators
- Google AI Edge Gallery: Interactive showcase app featuring two new demos powered by EmbeddingGemma 2
- Instant Media Search: Users find specific images or videos locally using natural language or example images
- Video Moments Finder: Locates specific visual moments across local video recordings without transcribing audio or generating intermediate text captions
- Google AI Edge Foresight (Mac): An experimental context-aware meeting companion that uses fully local AI processing for note-taking, indexing, and recalling conversations
- ML Kit: Android-focused integration for production apps with standardization and platform integration
- MediaPipe Tasks SDK: A unified developer experience spanning iOS, macOS, Windows, Linux, and Web
Available Demos & Tools
- Google AI Edge Gallery: Two new interactive showcases - Instant Media Search and Video Moments Finder
- LiteRT Benchmark Tool: Allows testing on various edge devices via Google Cloud
- GitHub Repository: Implements the model and provides guidance for building custom experiences
Contributors
Significant contributions were made by: Abhishek Jatram, Aditya Srivastava, Akshat Sharma, Alexander Kanaukou, Alice Zheng, Ami Kubota, Ander Dobo, Andrew Zhang, Brad Lassey, Chandramouli Amarnath, Chanchal Raj, Charlie Xu, Chenchen Tang, Chirag Gupta, Chintan Parikh, Cormac Brick, David Chou, Debapriya Maji, Denis Daletski, Fengwu Yao, Florian Kübler, Geonsun Lee, Gregory Karpiak, Himangshu Roy, Henrique Schechter, Vera, Hriday Chhabria, Ian Ballantyne, Ivan Llanos, Jae Yoo, Jenn Lee, Jianing Wei, Jing Jin, Jingjiang Li, Jingtao Zhou, Jingxiao Zheng, Juhyun Lee, Julius Kammerl, Karthik Thirumalai, Kat Black, Kristen Quan, Kristen Wright, Lu Wang, Lutz Justen, Malini PV, Marissa Ikonomidis, Matthew Chan, Matthew Soulanille, Matthias Grundmann, Mogan Shieh, Na Li, Naina Singla, Olivier Lacombe, Priya Patel, Qidong Zhao, Queenie Zhang, Renjie Wu, Rishika Sinha, Rishubh Khurana, Ronald Wotzlaw, Sachin Kotwani, Sandeep Patil, Sebastian Russo, Sebastian Schmidt, Shengyi Lin, Shuangfeng Li, Steven Toribio, Suril Shah, Sylvain Vignaud, Somdatta Banerjee, Tenghui Zhu, Vinod Mamillapalli, Vladimir Kirilyuk, Wai Hon Law, Weiyi Wang, Xiaoming Hu, Xinan Cheng, Xu Chen, Yi-Chun Kuo, Yishuang Pang, Yu-hui Chen
This reformatted article preserves all original facts, names, numbers, and quotations while organizing the content into clear, scannable sections for a technical audience.
Comments
No comments yet. Start the discussion.