Interestana
Home/News/Cloudflare AI Training Block Spares Googlebot
SEMrush Blog3 min read

By Interestana AI Editorial — AI-drafted, human-overseen. How we report

Cloudflare AI Training Block Spares Googlebot

Cloudflare announced an update to its "Disallow AI Training" setting, which now explicitly exempts Googlebot from being blocked from accessing websites. This change means that Google's web crawlers, including those used for AI training purposes, will continue to access content on sites that have enabled this specific Cloudflare setting. The company stated that this exemption is in recognition of Google's significant role in indexing the web and making information accessible. However, other major search engine crawlers, such as Microsoft's Bing, will not honor this exemption until early 2027. This staggered implementation highlights the ongoing complexities and negotiations between web infrastructure providers, search engines, and AI developers regarding data access for model training.

The "Disallow AI Training" setting was introduced by Cloudflare to provide website owners with greater control over how their content is used to train artificial intelligence models. Previously, enabling this setting would block a broad range of crawlers identified as being involved in AI training. The updated policy reflects a nuanced approach, acknowledging that not all web crawlers serve the same purpose or have the same impact on the web ecosystem. By exempting Googlebot, Cloudflare aims to balance the need for AI development with the concerns of website owners about unauthorized data scraping. This move is particularly significant given the vast amount of data Google indexes and the potential for this data to be incorporated into AI models.

The decision to exempt Googlebot while deferring Bing's compliance until 2027 suggests a strategic alignment with one of the dominant players in web search and AI development. It also implies that Cloudflare is engaging in ongoing discussions with various technology companies to refine its policies. The early 2027 timeline for Bing's adherence indicates that significant technical or policy coordination is required before Microsoft's crawlers will respect the AI training block exemption. This extended period allows for further development of AI training identification protocols and potential industry-wide standards.

Website owners who utilize Cloudflare's services can now have more confidence that their content will remain accessible to Google's indexing and AI training efforts, even if they have opted to block other AI training crawlers. This granular control is crucial for businesses and individuals who want to manage their digital footprint and ensure their data is used ethically and with appropriate consent. The evolving landscape of AI development necessitates such flexible tools, allowing for adaptation to new technologies and changing data governance requirements. Cloudflare's proactive adjustments to its settings demonstrate a commitment to serving its diverse customer base while navigating the rapidly advancing field of artificial intelligence.

Original source — read the full reporting at the publisher:

Read on SEMrush Blog

Get the weekly AI digest

AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.

Read next