- Google’s Googlebot crawler — long the backbone of its search index — has been quietly repurposed to simultaneously feed Gemini, its flagship AI model, creating a structural conflict of interest at the heart of the web’s economics: publishers cannot block AI data scraping without also blocking their search traffic, since both functions share the same bot; Cloudflare CEO Matthew Prince, whose firm handles traffic for roughly 20% of the global web, has documented that Google’s crawler sees 3.2x more of the web than OpenAI’s and 4.8x more than Microsoft’s — an advantage directly derived from Google’s 90% search market share and its ability to bundle AI scraping with search indexing that competitors cannot replicate.
- The economic threat to the open web is concrete and accelerating: Cloudflare data shows AI agents generated more than 57% of web traffic in 2026 — the first time in history that machine activity exceeded human activity on the internet; when AI systems read, summarize, and synthesize web content without sending human readers back to the source pages, the advertising and subscription revenue that incentivizes original content creation evaporates; internal Google slides disclosed in a US antitrust case show the company considered giving publishers an opt-out from AI data collection in April 2024 before launching AI Overviews, then decided not to because it was “evolving into a space for monetization” — a deliberate choice to exploit publisher dependence on Google search traffic to extract free AI training data.
- Two external forces are now forcing Google’s hand: Cloudflare issued an ultimatum that from September 15 it will block mixed-purpose crawlers by default for its ad-supported website customers — potentially cutting Google off from millions of sites — and the UK’s Competition and Markets Authority ordered in June that Google must give publishers a clear choice to block their content from AI products while remaining in search results, with an explicit prohibition on penalizing opt-out sites in search rankings; Google confirmed it is testing a global opt-out setting that lets sites block AI Overviews without affecting search rankings, and plans to roll it out worldwide after UK testing is complete.
- The opt-out fix addresses the symptom but not the structural problem: Google’s crawler remains technically unified, meaning sites still must trust Google to respect their opt-out preferences rather than being able to enforce exclusions themselves at the network level; Cloudflare’s Prince argues the only genuine solution is a full technical separation of Google’s crawler into two distinct bots — one for search indexing, one for AI scraping — so that publishers can block AI data collection directly and independently, without relying on Google’s self-reported compliance; the precedent matters because the same bundling dynamic is emerging across every major AI company that also operates a search or distribution product.
What Happened?
Bloomberg Opinion columnist Parmy Olson details how Google has been using its dominant search crawler (Googlebot) to simultaneously feed its Gemini AI model, making it structurally impossible for publishers to block AI data scraping without also losing search visibility — the traffic lifeline most websites depend on. Cloudflare, which handles traffic for ~20% of the web, has given Google a September 15 deadline to separate its crawlers or face a default block for millions of sites. The UK’s CMA separately ordered Google to provide a genuine opt-out from AI training without search ranking penalties. Google confirmed it is testing a global opt-out setting. AI agents now generate 57% of web traffic, exceeding human activity for the first time.
Why It Matters?
Google’s crawler bundling is the most important structural issue in the AI-and-media conflict: it lets Google extract AI training data from the entire web at no cost by leveraging a search monopoly that other AI companies cannot match, while simultaneously removing publishers’ ability to withhold consent without self-destructing. The internal slides showing Google deliberately chose not to offer an opt-out “because it was evolving into a space for monetization” confirm this was a deliberate policy choice, not an oversight. As AI agents displace human web traffic — now at 57% and rising — the economic model that funds original journalism, research, and content creation is being systematically dismantled.
What’s Next?
Watch the September 15 Cloudflare default-block deadline — whether Google separates its crawlers or finds a workaround will be the key near-term signal; watch whether Google’s global opt-out rollout includes meaningful enforcement mechanisms or relies on honor-system compliance; watch US antitrust proceedings against Google, where the internal 2024 opt-out slides are already in evidence and may inform remedies; and watch whether the EU and other major regulators follow the UK CMA’s lead in mandating genuine publisher control over AI data collection — the UK order could become the global template if it proves enforceable.
Source: Bloomberg Opinion














