Dispatch
OpenAI Announces $200B Valuation Round   •   EU AI Act Compliance Deadline Extended to 2027   •   Google DeepMind Releases Gemini Ultra 3.0   •   Y Combinator S26 Batch: 60% of Startups Are AI-Native   •   MarTech Consolidation: Salesforce Acquires MadTech Pioneer   •   LLM Token Costs Drop 80% Year-Over-Year   •   Meta Llama 4 Released Under Permissive Commercial Licence   •   Anthropic's Claude Achieves New Benchmarks on Reasoning Tasks   •   Venture Capital Flows to AI Infrastructure Exceed $4B in Q2   •   Adobe GenStudio Reaches 500,000 Enterprise Users   •   OpenAI Announces $200B Valuation Round   •   EU AI Act Compliance Deadline Extended to 2027   •   Google DeepMind Releases Gemini Ultra 3.0   •   Y Combinator S26 Batch: 60% of Startups Are AI-Native   •   MarTech Consolidation: Salesforce Acquires MadTech Pioneer   •   LLM Token Costs Drop 80% Year-Over-Year   •   Meta Llama 4 Released Under Permissive Commercial Licence   •   Anthropic's Claude Achieves New Benchmarks on Reasoning Tasks   •   Venture Capital Flows to AI Infrastructure Exceed $4B in Q2   •   Adobe GenStudio Reaches 500,000 Enterprise Users
Est. MMXXV — Independent Digital PressWednesday, 17 September 2026Vol. I — No. 204
MarTech • Startups • LLMs • Digital Strategyterekhindigital.comMorning Edition

Terekhin Digital Media

Rigorous Journalism at the Frontier of Digital Commerce & Machine Intelligence

Wednesday, 17 September 2026Issue No. 204
MarTech

Cloudflare Lets Sites Block AI Training Without Sacrificing Search Indexing

Launched 15 September: a 'Disallow AI Training' setting appends no-training preferences using Google-Extended and Applebot-Extended tokens while keeping Googlebot, Applebot, and Bingbot crawling for search. Solves the binary choice publishers have faced since 2024. Cloudflare's accompanying accountability framework requires crawler operators to honour opt-outs and guarantee training disallowance won't damage rankings.

Cloudflare launched a "Disallow AI Training" setting on 15 September that appends a no-training preference to robots.txt while explicitly allowing Googlebot, Applebot, and Bingbot to continue crawling for search indexing. The setting uses Google-Extended tokens for Gemini training opt-out and Applebot-Extended for Apple AI, with Bing support forthcoming. A separate "Block All" option halts all three major crawlers entirely for publishers who prefer complete disengagement. The mechanism solves the previously binary choice: publishers could either accept AI training on their content or remove themselves from search indexing. Cloudflare's accompanying accountability framework requires crawler operators to honour robots.txt training opt-outs, provide opt-out mechanisms for AI-generated summaries, and assure explicitly that training disallowance will not damage traditional search rankings. For publishers in AI licensing negotiations, the tool establishes a content-protection baseline that does not sacrifice SEO visibility — and a CDN-layer enforcement mechanism that does not require Google's cooperation to implement.

CloudflareAI trainingrobots.txtGEOpublisher toolsGoogle-ExtendedAI crawlerscontent protection
← Return to Front Page
Related Articles
© MMXXVI Terekhin Digital Media — All Rights Reserved — An Independent Digital Publication