robots.txt vs noindex: the distinction that causes the most damage

What robots.txt does, what noindex does, why they are not interchangeable, and the specific mistake that leaves unwanted pages permanently stuck in the index.

Published ·4 min read·1 sources cited

The short version

  • robots.txt controls crawling. noindex controls indexing. They are not alternatives and they do not substitute for one another.
  • A page blocked in robots.txt can still be indexed — and its noindex tag will never be seen, because the crawler cannot fetch the page.
  • To remove a page from the index: allow crawling and add noindex. Blocking is the opposite of what you want.
  • Use robots.txt to save crawl on things you never want fetched. Use noindex for things you do not want listed.

This is the most consequential misunderstanding in technical SEO, and it produces a failure that is both common and self-perpetuating: a page somebody wanted removed from search, blocked in robots.txt, still appearing in results months later with no way to shift it.

The mechanism is worth understanding precisely, because once you see it the correct procedure is obvious.

What each one actually does

Two different controls
Aspectrobots.txtnoindex
ControlsWhether a bot may fetch the URLWhether the page may be listed in results
Where it livesOne file at /robots.txtMeta tag in the page, or an HTTP header
Requires fetching the pageNoYes — the tag must be read
Prevents indexingNoYes
Prevents crawlingYesNo
Saves crawl budgetYesNo
Correct for private areasYes, alongside real authenticationAlso useful
Correct for removing a listed pageNo — actively harmfulYes

The third row is the whole problem. noindex is an instruction written on the page. If you have told the crawler it may not fetch the page, it never reads the instruction. The page stays in the index, usually with no description, and nothing you add to it will ever be seen.

The correct procedure for each goal

What to use when
GoalDo this
Remove a page from search resultsAllow crawling, add noindex, wait for recrawl
Stop bots wasting crawl on parameter URLsDisallow in robots.txt — these were never indexed
Keep a staging site out of searchHTTP authentication. Not robots.txt, not noindex alone
Remove a page urgentlynoindex plus the removals tool in Search Console
Prevent a PDF being indexedX-Robots-Tag: noindex HTTP header
Stop AI crawlers fetching your contentrobots.txt with their user agents — see the guide
noindex, the two ways to set it
<!-- In the page head -->
<meta name="robots" content="noindex, follow">

<!-- Or as an HTTP header, for non-HTML files such as PDFs -->
X-Robots-Tag: noindex

Note noindex, follow rather than noindex, nofollow. You generally still want links on the page to be followed so that authority flows through to pages you do want indexed. Using nofollow as well cuts that off for no benefit.

The staging site case

Neither control is the right answer for a staging environment. robots.txt is a request that well-behaved crawlers honour and nothing else does, and noindex requires the page to be fetched by anyone who asks. Neither prevents a human or a badly behaved bot from reading your unreleased site.

Use HTTP authentication. It is the only control that actually restricts access, and it has the useful side effect of making the staging site uncrawlable by definition. This is also how the single most damaging robots.txt accident happens — a staging Disallow: / promoted to production along with everything else.

How long removal actually takes

Adding noindex does not remove a page immediately. The crawler has to revisit the URL, read the instruction and then drop it, which can take days or weeks depending on how often that page is crawled. Low-value pages are crawled infrequently, which means the pages you most want gone are often the slowest to go.

If you need something out quickly, the removals tool in Search Console suppresses a URL from results temporarily while the noindex does its work. It is a stopgap rather than a fix — the suppression expires — so use both together: removals tool for speed, noindex for permanence. Resubmitting the affected URLs can also prompt a faster recrawl.

Bing and other engines have their own removal mechanisms and their own timelines, so a page gone from Google may persist elsewhere. If removal matters for a legal or reputational reason rather than a tidiness one, handle each engine separately through its webmaster tools — and remember that neither mechanism affects third-party archives or caches, which are outside your control entirely.

Check your production robots.txt right now

Visit /robots.txt on your live site and read it. A Disallow: / inherited from a staging configuration will remove your site from search results, and it is one of the most common catastrophic SEO incidents there is. Thirty seconds, and it is the highest-value check on this page.

Frequently asked questions

What is the difference between robots.txt and noindex?

robots.txt controls whether a crawler may fetch a URL; noindex controls whether a page may appear in results. Blocking in robots.txt does not remove a page from the index and prevents any noindex tag from ever being read.

Why is my blocked page still in Google?

Because blocking prevents crawling, not indexing. If other pages link to it, it can be indexed from those links alone — and the crawler cannot fetch it to see any removal instruction. Unblock it and add noindex.

Should I use both together?

Not on the same URL at the same time, and not in that order. Use noindex first, wait for the page to drop out of the index, then add a robots.txt block afterwards if you also want to save crawl.

How do I keep a staging site out of search?

HTTP authentication. robots.txt is only a request and noindex still requires the page to be publicly fetchable.

Sources

Every figure on this page traces to one of these. Dates are when we last read the page — pricing and features change, so treat anything older than a few months as a starting point rather than gospel. All outbound links here are nofollow.

  1. [1]
    Google Search Console

    Google · search.google.com · Vendor page · read 2026-09-18

Keep reading

This is the most consequential misunderstanding in technical SEO, and it produces a failure that is both common and self-perpetuating: a page somebody wanted removed from search, blocked in robots.tx…