How Box is unlocking multimodal enterprise agents with Gemini Embeddings 2
Box and Google Cloud are integrating multimodal embeddings into Box’s Agentic Platform using Gemini Multimodal Embeddings 2 to enable AI agents to process text, images, tables, and charts together for enterprise workflows.
Box’s Agentic Platform is expanding beyond text-based search to handle multimodal enterprise content, including financial tables, clinical images, and flowcharts, by integrating Google Cloud’s Gemini Multimodal Embeddings 2. This shift addresses limitations of traditional RAG systems, which struggle to preserve spatial relationships in structured documents like spreadsheets or interpret visual data such as medical scans. The update aims to mirror human-like document comprehension, ensuring column headers align with data points and visual elements are searchable alongside text.
The new capabilities enable cross-format retrieval, allowing AI agents to query images, charts, and text simultaneously within a unified semantic space. For example, users can locate a specific chart in a slide deck without manual tagging or verify a signed PDF contract against an email thread. The system supports formats like PDF, Excel, PowerPoint, PNG, and CSV while maintaining structural integrity, reducing the need for manual data alignment across disparate files.
In corporate finance and audit workflows, the multimodal embeddings enhance analysis of structured documents by preserving table layouts and visual trends. Financial teams can now align column headers with metrics, cross-reference written summaries with bar charts, and retrieve exact supporting data points instantly. This reduces errors in automated analysis where context is often lost in text-only indexing.
Healthcare teams benefit from the ability to synthesize visual and textual clinical data, such as linking patient photos to lab reports or triage grids. The system flags anomalies like rare parasitic patterns in microscopy images and cross-references findings with risk frameworks to provide immediate warnings. For legal and compliance teams, the technology audits visual documents against text records, identifying discrepancies such as outdated pricing in images or missing clauses in contracts.