AI Crawlers Now Cost More Than Real Users

The Linux kernel’s official Git server is spending more compute serving AI scrapers than it spends on every legitimate visitor combined. That’s the warning from Konstantin Ryabitsev, surfaced and amplified by Simon Willison, who flags it as a growing threat to anyone running a large, crawlable website. What stands out here is the scale: this isn’t a nuisance anymore. It’s the dominant workload.

Ryabitsev’s numbers, as reported by Simon Willison, are blunt. Across five geo-distributed nodes for git.kernel.org, 14 CPU cores are doing nothing at any given moment but rendering Git commits as HTML for bots. His summary: “we spend more CPU cycles rendering commits for scrapers than we spend on all other kinds of legitimate access, including git clones.”

Willison takes it personally, and for good reason. He builds Datasette, a tool that publishes data as huge numbers of crawlable web pages. “I worry about this a lot,” he writes. If your site exposes millions of URLs, you’re the exact target this problem hunts.

What’s actually changing

Crawling isn’t new. Search engines have indexed the web for decades. What’s different now is the economics behind the bots.

  • Volume exploded. The AI training and retrieval boom put dozens of new crawlers on the road, many ignoring robots.txt and hammering every link they find.
  • The cost landed on you. Rendering dynamic pages, like Git commits or database views, burns real CPU. The scraper pays nothing. You pay for the servers.
  • The traffic looks legitimate. These bots rotate IPs, fake user agents, and blend into normal request patterns, which makes simple blocking a losing game.

Ryabitsev calls it “background radiation.” That’s the right frame. It’s ambient, constant, and it degrades everything running underneath it.

Why it matters now

This is a tax on open infrastructure. Projects like kernel.org exist to serve developers freely. When most of the compute budget goes to feeding someone else’s model, the open web gets more expensive to run, and the people funding it start asking hard questions.

Look one to three years out and the pressure only builds. More AI companies, more retrieval agents pinging sites in real time, more autonomous crawlers. The volume curve points up. Sites that publish rich, dynamic content, the kind AI systems most want to eat, will feel it first and worst.

The likely response is a web that closes up. Expect more login walls, more aggressive rate limiting, more Cloudflare-style bot challenges, and more content that simply disappears behind APIs with paywalls. That’s a direct cost to the open, linkable web that made all of this possible in the first place.

What to do about it

If you run a large or dynamic site, treat crawler load as a first-class capacity problem, not an afterthought.

  1. Measure the split. Separate bot traffic from human traffic in your logs. You can’t manage what you haven’t quantified. Ryabitsev’s team clearly did.
  2. Cache aggressively. Pre-render and cache expensive dynamic pages so a scraper hit costs a file read, not a CPU-heavy render.
  3. Rate-limit by behavior. IP and user-agent blocking fail against rotating bots. Throttle on request patterns instead.
  4. Use a bot-management layer. Cloudflare, Fastly, and similar services now offer AI-crawler controls. For high-value dynamic content, they’ve moved from optional to necessary.
  5. Decide your policy on purpose. Choose what you let AI systems index, and enforce it. Hoping robots.txt gets respected is not a plan.

Willison’s track record on spotting infrastructure shifts early is strong, and he’s pointing at this one hard. The bill for the AI boom is landing in an unexpected place: the CPU budgets of the people who publish the open web. Plan your capacity as if the bots are your biggest users, because for a growing number of sites, they already are. Full details are at the original source.

Scroll to Top