Why Crawl Budget Shapes Crawler Behaviour
Search crawlers do not fetch everything they could. Each site receives a share of finite crawling capacity, and the way that share is calculated explains most of the behaviour publishers observe in their logs.
Capacity is finite and contested
A crawler operating across the web has fixed infrastructure and an effectively unlimited set of candidate pages, so every fetch spent on one site is a fetch not spent elsewhere.
Allocation therefore depends on expected value: how likely a fetch is to find content worth indexing, and how likely that content is to be requested by users.
Sites with large numbers of low-value pages consume their allocation on those pages, which is why crawl efficiency matters more on big sites than on small ones.
Server responsiveness sets the rate
Crawlers monitor response times and error rates and adjust their request rate accordingly, because overloading a site is both impolite and counterproductive.
Consistently fast responses lead to a higher rate, while slow responses or server errors cause the crawler to back off, sometimes sharply and for longer than the underlying problem lasted.
This creates a feedback loop where a period of instability reduces crawling for some time afterwards, and the reduction is not immediately reversed when performance recovers.
Revisit frequency follows observed change
Crawlers estimate how often each page changes and schedule revisits to match, so pages that change constantly are fetched frequently and static pages are fetched rarely.
The estimate is built from history, which means a page that has been stable for a long time is revisited slowly even after it changes substantially.
Freshness signals such as sitemap timestamps influence this, though they are treated as hints to be verified rather than as facts, because they are self-reported.
Duplicate and near-duplicate pages waste allocation
Sites that generate many similar pages through parameters, filters and sort orders present a large surface of pages with little distinct content.
Crawling those consumes allocation that would otherwise reach genuinely new content, and the crawler cannot tell which is which until it has fetched them.
Canonical declarations and parameter handling exist to solve this, and they work by reducing the candidate set rather than by increasing the allocation.
Budget explains behaviour that looks like a problem
Publishers frequently interpret irregular crawling as a penalty or a technical fault when it is ordinary allocation behaviour responding to observed value and responsiveness.
It also explains why blocking or throttling a verified crawler has lasting effects. The reduced rate persists after the block is lifted, because the crawler is working from a history that now includes a period of unavailability.