OFICIAL Google Cloud Blog

How KDDI built Buffmee, a faster, reliable consumer RAG app

What happened
Based on Google Cloud Blog · Sep 08, 2026

KDDI’s Buffmee consumer RAG app cut response latency by 38% and improved groundedness by 25% using Google Cloud’s automated evaluation and analytics tools, enabling faster, reliable AI-powered learning.

How KDDI built Buffmee, a faster, reliable consumer RAG app
Google Cloud Blog — Google
Key points
·
KDDI’s Buffmee RAG app reduced total response latency by 38% using Google Cloud’s automated evaluation and analytics tools.
·
The team improved groundedness scores by 25% by implementing LLM-as-a-Judge and the Rule of Hundreds in the evaluation framework.
·
KDDI cut evaluation workload by 75% through strategic content sampling across a two-dimensional difficulty grid of file formats and media composition.
Key numbers
·
KDDI improved groundedness scores by 25% and reduced total application response latency by 38%, hitting their target performance.
·
The development team transitioned to binary evaluation metrics, reduced evaluation workload by 75% through strategic content sampling, and calibrated thresholds based on product judgment.
·
KDDI’s Buffmee consumer RAG app cut response latency by 38% and improved groundedness by 25% using Google Cloud’s automated evaluation and analytics tools, enabling faster, reliable AI-powered learning.

KDDI, Japan’s major telecom carrier, launched Buffmee, a consumer-facing RAG app that grounds responses in over 100 sources including books and magazines. The service lets users search, summarize, and explore personalized learning while citing sources to ensure reliability. Engineers initially faced latency issues that prevented meeting target response times, prompting the need for a systematic approach to evaluation and performance optimization.

To address performance bottlenecks, KDDI implemented Google Cloud’s automated evaluation framework and performance optimization techniques. The team used the Gemini Enterprise Agent Platform Evaluation Service, including LLM-as-a-Judge and the Rule of Hundreds, to replace manual testing with a data-driven process. They ingested their document corpus, built hundreds of automated tests, and created a benchmark dataset to measure answer reliability across use cases.

KDDI improved groundedness scores by 25% and reduced total application response latency by 38%, hitting their target performance. The team analyzed production logs with BigQuery Agent Analytics and the ADK log analysis agent to visualize how prompt complexity impacted Time To First Token (TTFT). They optimized system prompts, reviewed sub-agent routing, and resolved deep-stack bottlenecks without sacrificing accuracy.

The development team transitioned to binary evaluation metrics, reduced evaluation workload by 75% through strategic content sampling, and calibrated thresholds based on product judgment. They also split massive system prompts exceeding 800 lines into modular ADK Skills to mitigate latency degradation and attention drift, enabling scalable, reliable generative AI applications.

Original source → Deals on Clipraptor.com →