llms.txt vs robots.txt: one is a standard, the other is a proposal

What robots.txt does, what llms.txt proposes to do, which one search engines and AI crawlers actually honour, and how to control AI crawler access properly.

Published ·3 min read·1 sources cited

The short version

  • robots.txt is a long-established standard with a published specification — RFC 9309 — that major crawlers honour.
  • llms.txt is a community proposal, not a standard. No major AI vendor has committed to honouring it as a ranking or access mechanism.
  • To control AI crawler access today, use robots.txt with the specific user agents. That is the mechanism that works.
  • Adding llms.txt is cheap and harmless. Treating it as an optimisation lever is not supported by anything published.

These two files get discussed as if they were alternatives at the same level of maturity. They are not. One is a decades-old protocol with an IETF specification that essentially every significant crawler implements. The other is a proposal that emerged from the community, has a reasonable rationale, and has not been adopted as a standard by the organisations whose behaviour it is meant to influence.

That distinction matters when you are deciding where to spend an afternoon, and it is worth stating plainly because a lot of content on this subject does not.

The two files compared

Status and function
Aspectrobots.txtllms.txt
StatusStandard — RFC 9309Community proposal
PurposeControl which crawlers may fetch which pathsOffer a curated, LLM-friendly guide to your site
Honoured by major crawlersYesNot committed to by major AI vendors
Affects accessYes — directlyNo
Affects citationYes — blocking removes youNo published evidence
Location/robots.txt/llms.txt
Cost to implementMinutesAn hour or two
Risk of getting it wrongHigh — you can deindex yourselfEssentially none

The risk row is the reason to be careful with the file that actually works. A misplaced Disallow: / in robots.txt can remove your site from search results, and it happens regularly — usually when a staging configuration reaches production. Nothing comparable can go wrong with a file nobody reads.

Controlling AI crawlers properly

If your goal is to allow or block AI crawlers, robots.txt is the mechanism. AI vendors publish the user-agent strings their crawlers use, and those are what you address. The important decision is directional: blocking them removes you from AI answers about your category rather than protecting anything you can monetise later.

Allowing AI crawlers while protecting private areas
User-agent: *
Allow: /
Disallow: /api/
Disallow: /dashboard

User-agent: GPTBot
Allow: /
Disallow: /api/
Disallow: /dashboard

User-agent: ClaudeBot
Allow: /
Disallow: /api/
Disallow: /dashboard

User-agent: PerplexityBot
Allow: /
Disallow: /api/
Disallow: /dashboard

Sitemap: https://example.com/sitemap.xml

Note that user-agent strings change and new crawlers appear. A blanket User-agent: * rule covers agents you have not enumerated, which is usually what you want — enumerate specific agents only when you want to treat one differently from the default.

Should you block AI crawlers?

For most businesses, no. Blocking means you are absent when an assistant answers a question about your category, and your competitors are not. The upside you are protecting — content not being used in training or summarisation — is real but rarely worth the cost of invisibility.

The exception is publishers whose entire business is monetised attention on the page. If a summary replaces the visit and the visit was the product, the calculation genuinely differs. That is a commercial judgement about your model, not an SEO decision, and it deserves to be made deliberately rather than inherited from a template.

If you do decide to publish an llms.txt, treat it as documentation rather than optimisation. A short, accurate description of what your site is and where the important sections live costs an hour and is harmless. What it will not do is influence whether you are cited, and building a workflow around maintaining it is effort better spent on the entity clarity and structured data that plausibly do.

Check what you are already blocking

A meaningful number of sites are blocking AI crawlers without knowing it, through a rule added years ago for an unrelated reason or inherited from a boilerplate configuration. Open your own /robots.txt and read it. It takes thirty seconds and it is the highest-value check on this page.

Frequently asked questions

Do I need an llms.txt file?

No. It is a community proposal that no major AI vendor has committed to honouring. Adding one is cheap and harmless; expecting it to affect citations is not supported by anything published.

How do I stop AI from using my content?

Block the relevant crawlers in robots.txt using their published user-agent strings. Be clear that this removes you from AI answers rather than simply protecting your content.

Does robots.txt stop indexing?

No — it controls crawling, not indexing. A blocked URL can still appear in results if other pages link to it. To prevent indexing use a noindex tag, which requires the page to be crawlable. See robots.txt vs noindex.

Should I block GPTBot?

Only if being absent from AI answers about your category is acceptable to you. For most businesses it is not; for some publishers it is a defensible commercial choice.

Sources

Every figure on this page traces to one of these. Dates are when we last read the page — pricing and features change, so treat anything older than a few months as a starting point rather than gospel. All outbound links here are nofollow.

  1. [1]
    RFC 9309: Robots Exclusion Protocol

    IETF · rfc-editor.org · Specification · read 2026-09-18

Keep reading

These two files get discussed as if they were alternatives at the same level of maturity. They are not. One is a decades-old protocol with an IETF specification that essentially every significant cra…