← Back to blog

About One Site in Twenty Blocks an Anthropic Crawler That Has Never Existed

The robots.txt lockout was counted twice, by teams who published their method. It was not a curve. It was a step that stopped at half.

Somewhere in the robots.txt files of about one website in twenty sits a line forbidding a web crawler that has never existed. The forbidden agent is called ANTHROPIC-AI, or CLAUDE-WEB. Anthropic owns neither and has never run either. On many of the same sites, the real Anthropic crawler, CLAUDEBOT, sits unblocked a few lines away. Measured across a large sample, a crawler that does not exist is disallowed more often than Meta's real one.

That single fact is the key to the whole story of the great robots.txt lockout, the wave of publishers who added Disallow lines for AI crawlers after OpenAI published the GPTBot user-agent on 7 August 2023. The lockout is real, it has been carefully counted, and almost everything intuitive about it is wrong. It was not a curve; it was a step that stopped at half. Its sector order is a monetization order in disguise. The fence it built is prospective only, around a field that had already been harvested. And the number itself measures something closer to what a site administrator had heard of in a given week than what the site actually consents to. The phantom crawler is where you can see that last part with your own eyes.

The curve already exists, so read it instead of drawing it

The temptation with a story like this is to go compute it yourself: pull the Wayback Machine's dated snapshots of robots.txt for the top domains, diff them month over month, and publish the adoption curve. Resist it, because the computation has already been run twice, by people who published their method, and the honest move is to read their numbers rather than re-derive them worse.

The Reuters Institute did the news sector in February 2024, pulling robots.txt for the fifteen most-used news sites in each of ten countries from the Wayback Machine for every available day of 2023. By the end of that year, 48 percent of those top news sites blocked OpenAI's crawlers and 24 percent blocked Google's, with the United States highest at 79 percent and Poland and Mexico near the bottom around 20 percent. Legacy print outlets blocked hardest, digital-born ones least. And not one site that blocked during 2023 unblocked again inside the year.

The Data Provenance Initiative did the whole web in July 2024, in a study of 14,000 domains sampled from the Wayback Machine at monthly intervals. Across the big training corpora, the share of tokens fully restricted by robots.txt rose from roughly 1 percent in mid-2023 to 5 to 7 percent by April 2024. In the head of the web, the largest and most actively maintained domains, it went from under 3 percent to somewhere between 20 and 33 percent in a single year, a relative jump the paper puts north of a thousand percent. The lockout is not a vibe. It is a measured, order-of-magnitude change in one year, and two independent teams using the same public archive agree on its shape.

It was a step, not a curve, and it stopped at half

Here is what the phrase "adoption curve" quietly hides. Ben Welsh runs a public tracker that pulls robots.txt from 1,152 news publishers twice a day. When Reuters cited it, in February 2024, about half of those publishers blocked OpenAI. The same tracker, same method, same population, on 12 September 2026, reports 49.7 percent. Thirty-one months, and the number moved by roughly nothing.

Everything the curve describes happened between August and December 2023. It is a step function: a five-month rise and then a thirty-one-month flat. Which means the interesting question was never how fast the defection spread. It is why it stopped at half. Somewhere around the turn of 2024, the population of publishers who were going to block had blocked, the population who were not had not, and the line went horizontal and stayed there for two and a half years.

Two honesty notes belong right here, because the flat line rests on them. The February 2024 figure is Reuters' prose summary of a chart, not a published table, so the comparison is approximate; but it survives a generous error bar, because even at "around 55 percent" in 2024 the move to 49.7 is small and possibly negative. And Welsh's population is roughly three-quarters US publishers, so this is a statement about mostly-American news, not the entire web. With those caveats, the shape holds: a wall went up fast and then stopped growing.

The flat line is hiding deals

A plateau sounds like stillness. This one is not. Underneath the net-zero, the largest publishers began moving in both directions at once, and the down-moves are purchases. The basis matters and is small: in a nine-domain sample chosen to span sectors, two had reversed. That is two cases, each verified by hand, and not a rate.

The Guardian's robots.txt in September 2024 named seven AI crawlers, including GPTBot and Google-Extended. By June 2026 it had removed exactly two of them: GPTBot and Google-Extended. Every other AI crawler on the list, including a real Anthropic agent and a fictional one, was still blocked. The Guardian announced a content partnership with OpenAI on 14 February 2025. Stack Overflow named GPTBot and Google-Extended in December 2023; by September 2024 both were gone while Bytespider stayed blocked, and it had announced an API partnership with OpenAI on 6 May 2024. Both deals are posted on OpenAI's own site, so the dates are not in doubt, and the robots.txt files were verified by reading them rather than trusting a keyword flag.

Neither site opened its doors. Each deleted the counterparty it had signed a contract with, deleted Google-Extended alongside it, and left the rest of the wall standing. Only OpenAI signed; nothing in the record explains the second removal, and two sites is a sample, not a pattern. That is what the flat line at half is concealing: not a stable consensus but a market clearing one publisher at a time. For an outlet large enough to be worth licensing, a robots.txt Disallow has stopped being a statement of principle and become a price tag. The block is the opening position; the deal is the block coming down for one buyer only.

The sector order is a monetization order wearing a sector's clothes

Which kinds of sites blocked, and in what order? The intuitive guess, news then e-commerce then forums, is wrong, and it is worth correcting because the true order explains itself. The measured sequence is news first, then forums and social media, then encyclopedias, with e-commerce at the bottom. In fact the Data Provenance Initiative concludes that crawls which honor robots.txt are set to shift toward e-commerce and organizational sites, precisely because the dynamic, user-generated sources are the ones walling themselves off. News tokens went from 3 percent fully restricted in 2023 to nearly 45 percent in the head sample; e-commerce barely moved.

The structural reason is in the same paper. The head of the web, news and social and encyclopedias, is about 73 percent of the heavily-trafficked domains, and it is far more monetized than the long tail of personal sites, blogs, and shops: more advertising, more paywalls. And the more monetized a site is, the more restricted it is, in both robots.txt and terms of service. The sites that blocked are the sites that were already selling their traffic. The order in which the sectors locked their doors is the order of how much they had to lose by leaving them open. Sector order is a monetization order wearing a sector's clothes.

One caution keeps this honest: these are token-weighted findings, about where the words are, not site-by-site timing. Amazon blocked GPTBot in the first month. No individual e-commerce site "came later." The claim is about the mass of the corpus, not any one shop's calendar.

What the file actually measures

Return now to the crawler that does not exist. The Data Provenance study found that 4.5 percent of websites disallow ANTHROPIC-AI or CLAUDE-WEB, agents Anthropic disclaims, while leaving the documented CLAUDEBOT unblocked, and that these phantom agents are restricted more often, about 62 percent of the time when any Anthropic agent is, than Meta's real crawler at 52 percent. There is only one way that happens. Someone published a copy-paste block list with a made-up agent in it, and it propagated across thousands of sites faster than anyone checked whether the agent was real.

That is the deep limit on every number in this essay, and it is better than any abstract warning about enforcement. A robots.txt disallow measures what a site administrator had heard of in the week they edited the file, not what the site consents to. The adoption curve is a curve of name recognition, with a consent curve somewhere inside it that no amount of careful counting can separate out. When a site blocks a crawler that was never crawling, the file is telling you about the administrator's information diet, not the publisher's will.

And the second half of the joke is the instrument itself. The only way to reconstruct this history is the Internet Archive's dated snapshots of robots.txt. The Internet Archive is disallowed by about a third of the sites that disallow anything. The lockout degrades the very tool used to measure the lockout. If you want one image for the whole affair, it is that: a fence built partly out of misremembered names, casting a shadow over the archive we have to squint through to see the fence.

The enclosure that enclosed nothing already taken

Read as a commons story, the lockout is a textbook defection: each publisher individually withdrawing from a shared training corpus that only had value collectively, and the curve is the shape of that withdrawal. But the fact that makes this more than a metaphor is that the fence is prospective only.

Common Crawl, the nonprofit whose scrape was about 60 percent of GPT-3's training data by the New York Times' own accounting, stores its archive in immutable WARC files, the format libraries and archivists use precisely because it cannot be quietly edited. Its executive director, Rich Skrenta, told The Atlantic in November 2025 that the format is immutable and nothing can be removed. The investigation reported that no archive files appeared modified since 2016, even as Common Crawl described its removal process for the New York Times and others as "50 percent, 70 percent, and then 80 percent complete." More than 900 news sites sit on the opt-out registry; a publishers' trade group sent a cease-and-desist in June 2026. Common Crawl's answer is that it filters the flagged URLs from future crawls and its public index, but it cannot un-write the WARCs already published.

So the defection curve is not a picture of the commons shrinking. It is a picture of new contributions stopping while the existing pool stays in circulation, immutable and already trained on. And that is exactly the property that makes each individual exit both rational and useless. Your Disallow protects nothing already in the corpus, and it costs nothing to add, so about half of the publishers with something to lose added it, half did not, and neither choice changed what the models had already read. The standard enclosure story says a fence went up around a shared field. The truer version is that the harvest was already in the barn, and the fences went up around the stubble, which is why so many people could not decide whether it was worth putting one up at all.

How to read these curves, and what a disallow is now

For anyone who deals with crawlers, robots.txt, or the confident percentages people quote about them, the usable lessons are three, and then one thing not to conclude.

First, a Disallow line is three different objects depending on who wrote it. For a publisher big enough to license, it is a price tag, the opening bid before a content deal drops it for one buyer. For nearly everyone else, it is name recognition, a copied list that may forbid crawlers that do not exist and miss the ones that do. And for the blanket blockers it is invisible to the tools that count: Reddit's file is a single group, User-agent: * then Disallow: /, which blocks every crawler on earth and names none of them, so any adoption curve that greps for the string "GPTBot" scores one of the most aggressive blockers on the web as wide open. Before you trust a lockout number, ask whether its author parsed robots.txt into groups or just pattern-matched the file, because the second one undercounts every Reddit.

Second, site-weighted and token-weighted curves answer different questions and will disagree, so always ask which denominator you are being handed. "45 percent of news tokens are restricted" and "49.7 percent of news sites block GPTBot" are both true and neither implies the other, because a handful of enormous sites carry most of the tokens. A percentage of the sites that expressed any preference is also not a percentage of the web: most of the web is silent, and silence is outside the measurement entirely, interpreted by the platforms and the AI companies to suit themselves.

Third, this method has a signature failure you must never mistake for a finding. A missing Wayback snapshot renders as an unblocking. Both published studies say so in their footnotes; the Reuters team names the countries where a gap in the archive faked a dip in the trendline. So when a raw curve shows a site reversing course, assume it is the archive's gap until you have read the actual file, because the real reversals are rare and, as we have seen, they are deals rather than changes of heart.

And the thing not to conclude from any of this is whether blocking works. Two vendor studies four months apart reached opposite answers on whether blocked sites still get cited by AI systems, neither was replicated, and both sell services adjacent to the question. The gap between what a robots.txt file says and what a crawler actually honors is real and, on the current evidence, unmeasured. It is a limit on this whole exercise, not a result of it.

The great lockout, counted, is a real number and a smaller, stranger one than the headline. It rose for five months and then held flat for thirty-one. It hides content deals underneath its plateau. It forbids crawlers that were never built and overlooks the blanket blockers that forbid everything. And it fenced a field that had already been harvested into an archive nobody can edit. The most honest sentence the data supports is the least dramatic one: about half of the publishers who had something to sell put up a sign, the sign means a different thing on every door, and the corpus it was meant to stop is immutable.

A line in robots.txt is a preference, not a provenance record

The lockout's deepest problem is that a Disallow line says what a site wants and records nothing about what was actually taken, by whom, or when — which is why a fence can go up around a field already harvested and nobody can tell from the fence. Chain of Consciousness is the other half: a verifiable log of what an agent saw, decided and asserted, so consent and use are both on the record instead of one being stated and the other inferred.

pip install chain-of-consciousness  ·  npm install chain-of-consciousness

Hosted Chain of Consciousness  ·  Verify a record

Sources

Figures note: the February 2024 news-blocking rate is a prose approximation of a chart rather than a published table, so the "moved by roughly nothing" comparison is approximate but survives a generous error bar; Welsh's population is about three-quarters US publishers, so the plateau is a claim about mostly-US news rather than the web. The phantom-agent and sector figures are the Data Provenance Initiative's. This essay reports only the adoption rate and its sector order; the enforcement gap between what a file says and what a crawler honors is named as a limit, because the available studies on it disagree and sell adjacent services, and it draws no conclusion about whether blocking changes what an AI system cites.