50,000 scrapes per reader: AI bots are draining publishers' content
Cloudflare has shown that some AI crawlers fetch a page as many as 50,000 times for every single visit by a reader. On top of that, 59 percent of Germans already use AI tools, but 81 percent want training data labelled.

Thesis: the ratio of scrapes to pageviews has stopped being an exchange and become an expropriation of the economics of content. Without licensing, the quality of information will turn into a cost nobody wants to bear.
On 1 July 2026 Cloudflare published its "Attribution Business Insights" overview. The scale of the problem is hard to ignore. The company set the number of page fetches by bots against the number of real reader visits. The ratios ran from 118:1 to nearly 50,000:1. In the most extreme cases, for every human who actually entered the site there are almost 50 thousand fetches by bots. Cloudflare splits traffic into training, search and agent categories. Only the second one brings the publisher any benefit.
The market is starting to react, because voluntariness is over. Cloudflare introduced rules for "mixed" bots that perform several functions at once. From 15 September they pass search traffic by default, but block training and agent activity on sites that carry ads. It was called the end of the free ride. Licensing deals are being signed at the same time. Netflix reached an agreement with Penske Media, the publisher of "Variety" and "Rolling Stone". Big players would rather pay for content than wait for regulation.
On the users' side the expectations are clear, and the ARD/ZDF Medienstudie 2026 survey quantifies them. 59 percent of Germans aged 14 and over use AI tools occasionally, and 37 percent do so weekly, compared with 22 percent a year earlier. Among the youngest, aged 14 to 29, 92 percent reach for them, and among the oldest, over 70, 15 percent. At the same time 54 percent approach model outputs with scepticism and check them for errors, while 81 percent consider transparent labelling of training data and sources important. Audiences do not want free magic. They want to know where the answer comes from.
There is a serious conflict of interest in this combination. The same users who want labels use tools that took content without a licence and without payment. Publishers, who were supposed to fund journalism from advertising, are losing revenue while supplying the material on which models are built that take over the role of intermediary in access to information. There is no symmetry here: the information broker does not bear the cost of producing it, yet still profits from it.
The solution is not to close off access but to put a price on it. Three elements are necessary: a public register of bots and their functions, a standard licensing mechanism available to small publishers too, and transparent indication of sources in model answers. Without the first we do not know who is fetching; without the second only the largest newsrooms will get deals; without the third a reader cannot tell synthesis from fact. Cloudflare showed that the data exists. Now it has to become the basis for an agreement, not for a complaint.
Sources
3- 01Cloudflare exposes AI crawlers hitting sites 50,000 times per visitorEN
- 02Declared good bots, mixed-use crawlers, gray scrapers: how AI accesses publisher contentEN
- 03ARD/ZDF-Medienstudie 2026: So nutzen die Deutschen KIDE
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.