The Discrepancy Between Scrapes And Referrals Actually Tells Us About The New Web

What a 1000 to 1 scrape ratio tells us about the future of content, the web's coming split, and where the profit really goes.

The Discrepancy Between Scrapes And Referrals Actually Tells Us About The New Web

AI companies scrape unprecedented amounts of content from publishers and return almost nothing in referrals, by some accounts, 1,000 scrapes for every 1 click. While this makes it immediately about the lost traffic, assuming AI is siphoning away clicks that would have otherwise gone to those sites, it's not quite that or rather, it's not just that.

In actuality, AI has decoupled two previously interconnected processes, now enabling someone to produce "valuable" information and someone else to commercially profit from it.

The way to understand this is to first look at what the web has been for the last 20 years or so, exposure → click → session → conversion. Impressions get logged in search console, sessions get logged in analytics, and somewhere down the line you try to attribute sessions to conversions. But this system is set up on the assumption that some user saw some page or domain that one could measure, which AI broke.

Now pages get scraped, folded into a model, and used to recommend something a user buys, without ever seeing the page. So can you still take traffic as a good proxy for influence? And does relative traffic between similar articles or domains really reflect their relative degree of influence?

Not anymore. An article with 20x the traffic of another might look like it has 20x the influence. But if the other's snippet gets shown to 200x as many users, it's actually the one with 200x more influence, regardless of what the dashboards say.

For a long time, marketing has had something called 'dark social', which is basically social media posts that get shared in private messages or some other way that the analytics software cannot follow. Hence, once a link gets clicked in a private message, the best one can do is bucket that traffic as direct.

With AI, there is now 'dark distribution': your articles get folded into a model, and get distributed to users in ways that you can no longer track and attribute.

Worse still, sometimes one might not even be cited at all. Let's say, your article contains a fact and someone drops that fact into their own piece, sitting alongside a bunch of others in a comparison roundup. The model might end up crediting them for it, not you, just because it's easier to appropriate than to create, aka 'citation arbitrage'.

For businesses, that means your content could show up a million times in AI responses and still, none of those users visit your site, none of them buy anything.

That's precisely the problem with the scrape-to-referral ratio: influence you can measure, still failing to translate into value, which is, according to me, a much harder thing to solve.

It is possible that we can see some companies pull back from making their best content freely available, or at least from making it easy to scrape and repurpose. Raw data can go behind a login, while only static write-ups of the findings stay public.

These are all going to be ways of tying content to value, now that AI limits how much value it generates for free.

The web will split into two categories if enough businesses start acting this way: one for information that most people are comfortable using for free and contributing to, for reputation, where there's no need to visit a particular domain because ANY knowledge repository is equivalent. And one for information that is valuable enough to be embedded in a product, where one can extract value from its use, and for which one has some audience that one can sell access to.

The scrape-to-referral ratio you're seeing might be less a harbinger of doom for traffic, and more a sign that we're entering a future where valuable knowledge only lives behind walled gardens, which, for content creators, is a far more interesting problem to have.

Sure, it raises some interesting questions about measurement philosophy, but isn't it also pointing at how valuable knowledge now travels without its creator attached, and what that does to the company that paid to create it in the first place?