robots-txt

2 posts

cloudflare

Content Independence Day, one year on- building the business model for the agentic Internet (opens in new tab)

Cloudflare argues that generative AI has rapidly replaced the traditional web model in which publishers traded content access for search referrals. With AI now driving much of online discovery and crawler activity, content is increasingly consumed without users visiting its source. The company says a new market is emerging in which transparency, access controls, scarcity, and licensing can help publishers regain economic value. ## AI’s rapid transformation of the Internet - Generative AI adoption has reached more than 2.5 billion regular users—over 30% of humanity—in roughly 3.5 years, reportedly more than twice the adoption speed of smartphones. - Users now spend only about 15 minutes on the open web for every hour spent searching for information. - Instead of visiting and comparing multiple websites, users increasingly receive consolidated answers directly from AI systems. - More than 50% of Internet traffic is now non-human, marking the arrival of what Cloudflare calls the “agentic Internet.” ## Crawlers are increasingly focused on AI - AI training accounted for 52% of crawler requests in June 2026, up from 22% in spring 2025. - Mixed-use crawlers, combining search, agent activity, and training, represented more than 36% of crawler traffic. - Traditional search crawlers make up a smaller share of activity, even though they remain important for sending visitors to publishers. - Mixed-purpose crawling makes it difficult for site owners to remain visible to AI-driven discovery without also giving away content for training without compensation. ## The traditional web business model is breaking down - Historically, publishers allowed search engines to crawl their content in exchange for visibility and referral traffic. - AI systems now answer questions, conduct research, compare products, and complete tasks without necessarily sending users to original sources. - Content can therefore be crawled, indexed, and monetized by AI companies while the original publisher receives little or no traffic. - News and media organizations experienced the disruption first, but retail, software, IT, finance, and other sectors are also affected. - Some heavily crawled categories have seen human traffic fall by as much as 40% in under a year. - Publishers are preparing for “Google Zero,” in which search referrals provide little meaningful traffic. ## The impact extends across industries - Any organization publishing proprietary information online may need a strategy for AI access and monetization. - The issue affects not only traditional publishers but also businesses whose websites contain valuable product, technical, financial, or industry knowledge. - Cloudflare frames the sustainability of online content as an economic and public-interest concern because the Internet remains a major global information resource. ## Building a market for content Cloudflare says Content Independence Day focused on three goals: - Give site owners transparency and control over how their content is accessed and monetized. - Create scarcity by allowing publishers to restrict or selectively permit AI access. - Establish a marketplace where publishers and AI companies can discover, license, and price content. According to the post, these efforts have helped create the early conditions for a monetized content market. ## Control and data create negotiating power - Cloudflare’s attribution, business intelligence, and enforcement tools let publishers observe AI access at the network level. - These tools provide stronger practical enforcement than voluntary mechanisms such as `robots.txt`. - Publishers can identify: - How often LLMs attempt to access their content - Which competing AI systems are crawling their sites - Which URLs are most in demand - The relationship between crawling and referrals - Restricting or controlling access creates scarcity, which gives publishers leverage in licensing negotiations. - Better operational data reduces information asymmetry and allows content owners to negotiate with evidence rather than guesswork. Ultimately, the post recommends treating online content as an economic asset rather than an unlimited free input. Publishers should measure AI consumption, control access, and pursue licensing arrangements so that the agentic Internet can support content creation instead of undermining it.

cloudflare

Google’s AI advantage: why crawler separation is the only path to a fair Internet (opens in new tab)

Google’s dominance in search gives it a structural advantage in generative AI: publishers must allow Googlebot to preserve search visibility, while Google can also reuse that access for AI products. The authors argue that this blurs search indexing and AI data collection, deprives publishers of traffic and compensation, and disadvantages competing AI companies. They support the CMA’s proposed UK conduct rules but say the only fair solution is to separate crawling for search from crawling for generative and agentic AI. ## CMA’s Strategic Market Status designation - The UK’s Digital Markets, Competition and Consumers Act 2024 allows the CMA to designate firms with substantial, entrenched market power as having Strategic Market Status. - In October 2025, Google received this designation for general search and search advertising, where it holds roughly 90% of the UK market. - The designation covers AI Overviews and AI Mode, allowing the CMA to impose legally enforceable conduct requirements on Google’s search ecosystem. - The authors view the CMA’s consultation as an important first step toward clearer rules for AI crawling and publisher control. ## Problems with Google’s dual-purpose crawler - Publishers cannot realistically block Googlebot because doing so could reduce their visibility in Google Search and damage advertising revenue. - Google uses the same search access not only for indexing and referrals, but also to ground AI Overviews, AI Mode, and broader generative AI services. - These AI features may reproduce publisher content while sending little or no traffic back to the original sites. - This threatens ad-supported publishing models and can put Google in direct competition with the publishers whose content it uses. - Unlike other AI companies, Google can obtain large amounts of content without negotiating payment, because publishers are effectively unable to refuse its search crawler. ## Google’s crawling advantage Cloudflare’s data indicates that Googlebot accesses substantially more unique pages than other major AI crawlers: - About 1.7 times more than ClaudeBot and GPTBot. - About 3 times more than Meta-ExternalAgent. - About 3.3 times more than Bingbot. - About 5.1 times more than Amazonbot. - Nearly 15 times more than Applebot. - Nearly 167 times more than PerplexityBot. - More than 700 times more than CCBot. - More than 1,800 times more than archive.org_bot. - Googlebot crawled roughly 8% of the sampled unique URLs during the two-month observation period. ## Limits of robots.txt and the need for separate controls - Publishers are much less likely to block Googlebot in `robots.txt` because of its importance for search referrals. - `robots.txt` expresses preferences but does not technically enforce crawler behavior; publishers must rely on bots to comply. - Web Application Firewalls can technically block unwanted crawlers, but this does not solve the core problem when search and AI access are tied to the same Googlebot identity. - The authors therefore argue that publishers need a meaningful, independent way to permit Google Search indexing while refusing the use of their content for generative AI. The proposed CMA rules should go further by requiring effective separation between search crawling and AI crawling. Publishers should be able to opt out of generative AI use without sacrificing search visibility, creating fairer conditions for content creators and competing AI developers.