Crawl budget vs index bloat: two problems people keep confusing

What crawl budget actually means, what index bloat is, why most sites have neither problem, and how to tell if yours genuinely does.

Published ·4 min read·3 sources cited

The short version

  • Crawl budget is about bots not reaching your pages. Index bloat is about low-value pages being indexed that should not be.
  • Most sites under a few thousand URLs have neither problem, despite a great deal of content suggesting otherwise.
  • Both are usually symptoms of the same cause: URL generation you did not intend — parameters, facets, filters, pagination.
  • The fixes are different and mixing them up causes damage: blocking in robots.txt does not remove pages from the index.

These two terms get used interchangeably and they describe opposite failure modes. Getting them confused leads directly to a common and damaging mistake: blocking pages in robots.txt to remove them from the index, which does not work and can make things worse.

It is also worth saying plainly that most sites reading about this have neither problem. Crawl budget is a genuine constraint on sites with hundreds of thousands of URLs. On a 400-page site it is not a thing you need to manage.

The two problems, distinguished

Different symptoms, different fixes
AspectCrawl budgetIndex bloat
The problemBots do not reach pages you want crawledPages you do not want are indexed
SymptomNew content takes a long time to appearThousands of indexed URLs you did not intend
Where you see itLog files, Search Console crawl statsSearch Console page indexing report
AffectsVery large sitesAny site generating URLs automatically
Usual causeWasted crawl on junk URLsJunk URLs being indexable in the first place
Correct fixReduce junk, improve internal linkingnoindex, canonical, or removal
Wrong fixBlocking everything in robots.txtBlocking in robots.txt — it does not deindex

The last row is the important one. robots.txt controls crawling, not indexing. A URL blocked there can still appear in results if something links to it — and because the crawler cannot fetch it, it also cannot see a noindex tag you may have added. Blocking a page you want removed is the one action guaranteed to keep it indexed indefinitely.

Do you actually have either?

  1. Open the page indexing report in Search Console.
  2. Compare indexed pages against the number of pages you believe you have. A large discrepancy upward is index bloat.
  3. Check how long new content takes to be indexed. Weeks on a small site points at something else; on a very large site it may be crawl budget.
  4. Crawl your site with Screaming Frog and count URL variants — parameters, filters, sort orders, pagination. That count is usually the whole story.
  5. If your site is under a few thousand URLs and nothing above looks wrong, you have neither problem. Go and work on content.

The shared root cause

Both problems almost always trace to unintentional URL generation. Faceted navigation producing a unique URL for every combination of filters. Session or tracking parameters creating duplicates of every page. Sort orders, pagination, print views, calendar pages extending infinitely into the future.

A 500-product catalogue with five filters can generate hundreds of thousands of crawlable URLs. That single architectural decision produces crawl waste and index bloat simultaneously, which is why the two get confused — they frequently arrive together from the same source.

The durable fix is upstream of both problems: stop generating the URLs. Faceted navigation that produces links for every filter combination, session parameters appended to every link, and calendars that page infinitely into the future are all architectural decisions that can be changed. Blocking or noindexing the output treats the symptom and leaves the generator running, which is why these problems tend to return a year later with a different parameter name.

The correct order of operations

To remove indexed pages: allow crawling, add noindex (or a correct canonical), wait for recrawl and deindexing, and only then consider blocking in robots.txt if you also want to save crawl. Blocking first strands the pages in the index permanently because the crawler can never see the instruction to remove them.

Frequently asked questions

What is crawl budget?

Roughly, how much crawling a search engine will do on your site in a given period. It becomes a genuine constraint on very large sites and is essentially irrelevant below a few thousand URLs.

What is index bloat?

Having large numbers of low-value pages indexed that you never intended to publish as destinations — parameter variants, filter combinations, pagination, internal search results.

Does robots.txt remove pages from Google?

No. It prevents crawling, not indexing. A blocked URL can still be indexed from links, and because it cannot be fetched, any noindex tag on it will never be seen. See robots.txt vs noindex.

How do I fix index bloat?

Allow crawling, apply noindex or a correct canonical to the pages concerned, and wait for recrawl. Better still, stop generating the URLs in the first place.

Should I worry about crawl budget on a small site?

No. Below a few thousand URLs it is not a meaningful constraint, and time spent on it is time not spent on content.

Sources

Every figure on this page traces to one of these. Dates are when we last read the page — pricing and features change, so treat anything older than a few months as a starting point rather than gospel. All outbound links here are nofollow.

  1. [1]
    Introduction to robots.txt

    Google Search Central · developers.google.com · Official documentation · read 2026-09-18

  2. [2]
    Google Search Console

    Google · search.google.com · Vendor page · read 2026-09-18

  3. [3]
    Screaming Frog SEO Spider Pricing

    Screaming Frog · screamingfrog.co.uk · Vendor page · read 2026-09-18

Keep reading

These two terms get used interchangeably and they describe opposite failure modes. Getting them confused leads directly to a common and damaging mistake: blocking pages in robots.txt to remove them f…