Securing internal systems against increasingly capable and imperfectly aligned AI — Google DeepMind
Google DeepMind introduces an AI Control Roadmap to secure internal AI agents, addressing risks from misalignment or unexpected behavior despite advanced capabilities.
Video
Video available
The useful question is what changes for users, developers or buyers, and whether the announcement stays industry context or becomes something people can actually use.
Google DeepMind has developed an AI Control Roadmap to manage risks posed by increasingly capable AI agents operating within its systems. The framework emphasizes a defense-in-depth approach, combining traditional safeguards like sandboxing and prompt injection resistance with model alignment as a primary defense layer. This ensures security even if AI alignment is imperfect, treating agents as potential misaligned entities requiring controlled access and oversight. The initiative reflects a cautious approach to deploying advanced AI while maintaining robust security measures.
The roadmap introduces a novel threat-modelling framework for AI, treating untrusted agents as potential insider threats. It adapts the MITRE ATT&CK framework to break down risks into smaller tactics and techniques, enabling methodical tracking and proactive defense testing. The system measures performance using coverage, recall, and time-to-response metrics to ensure effectiveness. Security protocols are mapped to measurable milestones based on AI capability levels, with detection and prevention levels tailored to model capabilities.
Google DeepMind has built an internal prototype of an asynchronous monitoring system to review AI agent activities, such as coding tasks, and escalate findings for human review. Analyzing over a million agent tasks has provided insights to refine safety protocols, moving beyond keyword filtering to identify behavioral patterns. This research informed the development of a live monitor for the Gemini Spark agent, enabling real-time responses to issues like unintentional data deletion, though most flagged events stem from misinterpretation rather than adversarial intent.
The AI Control Roadmap is part of a multilayered approach to agent security, integrating protocols with broader infrastructure to safely manage advanced AI models. Google DeepMind also published a technical framework, 'Three Layers of Agent Security,' aimed at policymakers to improve security at individual agent, multi-agent system, and ecosystem levels. The goal is to foster collaboration among industry, policymakers, and academia to build societal resilience and confidently deploy capable AI while establishing a secure foundation for the future.