OFICIAL Google Cloud Blog

How Box is unlocking multimodal enterprise agents with Gemini Embeddings 2

What happened
Based on Google Cloud Blog · Aug 18, 2026

Box and Google Cloud are integrating multimodal embeddings into Box’s Agentic Platform using Gemini Multimodal Embeddings 2 to enable AI agents to process text, images, tables, and charts together for enterprise workflows.

How Box is unlocking multimodal enterprise agents with Gemini Embeddings 2
Google Cloud Blog — Google
Key points
·
Enterprise content management is experiencing its biggest architectural shift since the cloud migration era.
·
For years, enterprises have stored trillions of gigabytes of critical data in Box: financial models, clinical trial protocols, M&A due diligence rooms, engineering schematics, and legal compliance playbooks.
·
Up to this point, text-based search and retrieval-augmented generation (RAG) have successfully unlocked the vast narrative knowledge within these repositories, establishing a powerful and highly effective baseline for enterprise AI intelligence.
·
Traditional RAG architectures have mastered text processing, but the agentic era demands more.

Box’s Agentic Platform is expanding beyond text-based search to handle multimodal enterprise content, including financial tables, clinical images, and flowcharts, by integrating Google Cloud’s Gemini Multimodal Embeddings 2. This shift addresses limitations of traditional RAG systems, which struggle to preserve spatial relationships in structured documents like spreadsheets or interpret visual data such as medical scans. The update aims to mirror human-like document comprehension, ensuring column headers align with data points and visual elements are searchable alongside text.

The new capabilities enable cross-format retrieval, allowing AI agents to query images, charts, and text simultaneously within a unified semantic space. For example, users can locate a specific chart in a slide deck without manual tagging or verify a signed PDF contract against an email thread. The system supports formats like PDF, Excel, PowerPoint, PNG, and CSV while maintaining structural integrity, reducing the need for manual data alignment across disparate files.

In corporate finance and audit workflows, the multimodal embeddings enhance analysis of structured documents by preserving table layouts and visual trends. Financial teams can now align column headers with metrics, cross-reference written summaries with bar charts, and retrieve exact supporting data points instantly. This reduces errors in automated analysis where context is often lost in text-only indexing.

Healthcare teams benefit from the ability to synthesize visual and textual clinical data, such as linking patient photos to lab reports or triage grids. The system flags anomalies like rare parasitic patterns in microscopy images and cross-references findings with risk frameworks to provide immediate warnings. For legal and compliance teams, the technology audits visual documents against text records, identifying discrepancies such as outdated pricing in images or missing clauses in contracts.

Original source → Deals on Clipraptor.com →