The short version
- Crawl budget is about bots not reaching your pages. Index bloat is about low-value pages being indexed that should not be.
- Most sites under a few thousand URLs have neither problem, despite a great deal of content suggesting otherwise.
- Both are usually symptoms of the same cause: URL generation you did not intend — parameters, facets, filters, pagination.
- The fixes are different and mixing them up causes damage: blocking in robots.txt does not remove pages from the index.
These two terms get used interchangeably and they describe opposite failure modes. Getting them confused leads directly to a common and damaging mistake: blocking pages in robots.txt to remove them from the index, which does not work and can make things worse.
It is also worth saying plainly that most sites reading about this have neither problem. Crawl budget is a genuine constraint on sites with hundreds of thousands of URLs. On a 400-page site it is not a thing you need to manage.
The two problems, distinguished
The last row is the important one. robots.txt controls crawling, not indexing. A URL blocked there can still appear in results if something links to it — and because the crawler cannot fetch it, it also cannot see a noindex tag you may have added. Blocking a page you want removed is the one action guaranteed to keep it indexed indefinitely.
Do you actually have either?
- Open the page indexing report in Search Console.
- Compare indexed pages against the number of pages you believe you have. A large discrepancy upward is index bloat.
- Check how long new content takes to be indexed. Weeks on a small site points at something else; on a very large site it may be crawl budget.
- Crawl your site with Screaming Frog and count URL variants — parameters, filters, sort orders, pagination. That count is usually the whole story.
- If your site is under a few thousand URLs and nothing above looks wrong, you have neither problem. Go and work on content.
The shared root cause
Both problems almost always trace to unintentional URL generation. Faceted navigation producing a unique URL for every combination of filters. Session or tracking parameters creating duplicates of every page. Sort orders, pagination, print views, calendar pages extending infinitely into the future.
A 500-product catalogue with five filters can generate hundreds of thousands of crawlable URLs. That single architectural decision produces crawl waste and index bloat simultaneously, which is why the two get confused — they frequently arrive together from the same source.
The durable fix is upstream of both problems: stop generating the URLs. Faceted navigation that produces links for every filter combination, session parameters appended to every link, and calendars that page infinitely into the future are all architectural decisions that can be changed. Blocking or noindexing the output treats the symptom and leaves the generator running, which is why these problems tend to return a year later with a different parameter name.
The correct order of operations
To remove indexed pages: allow crawling, add noindex (or a correct canonical), wait for recrawl and deindexing, and only then consider blocking in robots.txt if you also want to save crawl. Blocking first strands the pages in the index permanently because the crawler can never see the instruction to remove them.