Have it both ways: stay discoverable in search while disallowing AI training
Cloudflare introduces a Disallow AI Training setting to let site owners block AI training while remaining search-indexed, with commitments from Apple, Google, and Microsoft to honor the preference.
Website owners previously faced an either/or choice between allowing AI training or maintaining search discoverability due to mixed-use crawlers performing both tasks. Cloudflare’s new Disallow AI Training setting resolves this tradeoff by publishing a robots.txt directive that blocks AI training while preserving search indexing. The setting is designed to work alongside existing controls, offering granular control over how content is used by crawlers. Operators including Apple, Google, and Microsoft have committed to respecting this preference, addressing a longstanding frustration for publishers.
The announcement builds on Cloudflare’s existing bot classification system, which distinguishes between crawlers by behavior rather than identity. Mixed-use crawlers—those performing both search and AI training—were previously difficult to manage without sacrificing discoverability. The new setting allows site owners to block AI training specifically, while still permitting search crawlers to index their content. This change applies to all training crawlers, including those from major tech companies, and aligns with Cloudflare’s broader goal of giving publishers more control over content usage.
Cloudflare also plans to introduce finer controls for AI summaries by early next year, allowing site owners to specify how much of their content appears in AI-generated summaries. This follows the company’s existing requirement for mixed-use crawler operators to provide an opt-out for AI summaries. The move reflects a broader shift toward transparency and publisher choice in how content is consumed by automated systems. Site owners will be able to manage these settings at the domain level, simplifying compliance with evolving standards.
Existing site configurations will automatically migrate to the new system, with legacy Block AI Bots settings converting to Disallow AI Training where applicable. New domains will be offered preset configurations based on whether they rely on advertising revenue, with more restrictive defaults for ad-supported sites. Cloudflare’s Accountable designation recognizes operators that meet transparency and control requirements, including Apple, Google, Microsoft, Amazon, Anthropic, Meta, and OpenAI, whose crawlers are already classified accordingly.