Wikimedia’s engineering team says AI-powered crawlers generate 65 % of the platform’s most expensive traffic, turning a harmless-looking bot surge into a line-item on server bills. The same pattern now shows up on tiny blogs and hobby sites.
Why the spike matters
Human visitors linger on popular pages—homepages, featured articles, and other well-known content. CDNs cache those pages at edge locations, so the origin server rarely does heavy lifting. AI crawlers, by contrast, scrape thousands of obscure URLs in a single run. Those pages are rarely cached, so every request travels all the way back to Wikimedia’s origin infrastructure. The result: a disproportionate amount of compute, storage and network usage for a relatively small slice of total traffic.
The engineering blog that released the data notes that while bots make up only a modest fraction of overall requests, they dominate the “costly” bucket.
Who feels the pain
Smaller operators lack that cushion. The hidden cost isn’t just money; it also risks degraded performance for legitimate visitors.
What can site owners do
- Start with rate-limiting – throttle the number of requests a single IP can make in a short window. This slows aggressive crawlers without outright denying them access.
- Weigh visibility against cost – blocking a search-engine bot may shave off traffic, but it can also drop your pages from indexes, hurting organic discoverability.
- Classify the bots – not all crawlers are created equal. Some belong to legitimate search tools that bring traffic; others are pure scrapers that add no value.
- Deploy proof-of-work challenges – require a small computational puzzle before serving a page. Human browsers solve it instantly; bots must expend extra cycles, raising their cost.
- Use CDN analytics – most CDNs surface per-bot request metrics. Identify which agents hit the origin most often and tailor rules accordingly.
The takeaway is simple: AI crawlers are no longer a free-riding curiosity; they’re a measurable cost center. Whether you run a global encyclopedia or a single-page portfolio, treating bot traffic as a line item in your infrastructure budget can prevent surprise bills and keep your site responsive for real users.
