How AI assistants pick sources — what is known, and what is guesswork

What can actually be established about how AI assistants choose which sites to cite, what is reasonable inference, and what is confident speculation sold as fact.

Published ·4 min read·2 sources cited

The short version

  • No major assistant publishes its source-selection criteria. Anyone presenting a ranked list of "AI ranking factors" is inferring, not reporting.
  • What is well established: assistants with live retrieval mostly draw on indexed, crawlable web content. Being findable remains the prerequisite.
  • Reasonable inference from observation: clearly stated, attributable claims from recognisable sources get quoted more than vague ones from unknown sites.
  • Treat this whole area as an engineering problem with an unknown spec. Optimise for being unambiguous and verifiable, because those help under every plausible mechanism.

This is the question everyone in the category wants answered and nobody can answer properly. The honest starting position is that the selection process is not documented, varies between products, changes without notice, and is not directly observable from outside. That has not stopped a substantial industry from publishing confident factor lists.

What follows separates three things that usually get blended: what is established, what is reasonable inference, and what is speculation. The practical advice at the end holds under all of them, which is the main reason to trust it.

What is actually established

  • Assistants with browsing or retrieval use web content. They fetch pages, and pages they cannot fetch cannot be used.
  • Crawler access is controllable by you. Your robots.txt governs which agents may fetch your pages, and blocking them removes you from consideration.
  • Some assistants surface citations with links. Which means the source set is at least partially inspectable for a given answer.
  • Training data and live retrieval are different mechanisms. A model may "know" about you from training without ever fetching your site, and that knowledge has a cutoff you cannot influence.

That last distinction matters more than it gets credit for. If an assistant describes your company from training data, no amount of on-site optimisation changes the answer until the next model. If it retrieves live, your current pages matter immediately. Most products now do some of both, which is why answers about you can be simultaneously current and out of date.

What is reasonable inference

Observed patterns and how much weight to put on them
PatternConfidenceBasis
Cited pages tend to already rank wellHighConsistently observed; retrieval commonly uses a search index
Clearly stated facts get quoted more than implied onesMedium-highExtraction favours self-contained statements
Recognisable, established sources are preferredMedium-highConsistent with how these systems are evaluated
Consistency across the web increases confidenceMediumPlausible mechanism; widely observed
Valid structured data helps extractionMediumStrong rationale, no published confirmation
Recency matters on fast-moving topicsMediumObserved, varies by product
A specific word count is optimalVery lowNo evidence whatsoever
A specific "AI readability score" mattersVery lowVendor-invented metric

Confidence ratings are our editorial judgement from observation and mechanism, not measured findings. We include the bottom two rows because they circulate widely and deserve naming.

What is speculation sold as fact

A number of claims circulate with a confidence their evidence does not support: that assistants weight particular schema types heavily, that an llms.txt file materially affects citation, that there is an optimal answer length, or that a proprietary score predicts citation likelihood. None of these has published support from the people who build these systems.

Some may well turn out to be true. The problem is not the hypotheses, it is presenting them as findings — particularly when the presenter sells a product that measures the thing they have declared important. Apply the same scepticism you would to any vendor-supplied ranking factor list. See llms.txt vs robots.txt.

What to do given the uncertainty

The useful move when a specification is unknown is to optimise for properties that help under every plausible mechanism. Fortunately those are also just good practice, which means the downside of being wrong is zero.

  1. Be retrievable. Check your robots.txt permits the agents you want, and that pages render without requiring JavaScript execution to produce their main content.
  2. Rank well. Under every observed pattern, ranking correlates with being in the candidate set. Ordinary SEO is not superseded here.
  3. State facts explicitly, with attribution and dates. "X costs $129/mo as of September 2026, per the vendor pricing page" survives extraction; "X is competitively priced" does not.
  4. Describe yourself in one plain sentence and use it consistently everywhere you appear.
  5. Keep structured data valid, checked in the Rich Results Test. The rationale is strong even without published confirmation, and it costs nothing.
  6. Be corroborated. Being described the same way in several credible places is a plausible confidence signal under any reasonable design.

Applying this standard to ourselves

This page could have been a ranked list of twelve AI ranking factors with confident percentages. It would rank better and it would be fabricated. We have instead said what is known, what is inferred and what is invented — which is the same standard we apply to the pricing tables elsewhere on this blog, and the reason any of it is worth reading.

Frequently asked questions

How do AI assistants choose which sites to cite?

No major assistant publishes its criteria. Observation suggests cited pages usually rank well already, come from recognisable sources, and state relevant facts clearly enough to extract. Anything more specific than that is inference.

Is there an AI ranking factor list?

Not a real one. Published lists are inferred from observation, often by companies selling tools that measure the factors they have declared important. Treat them accordingly.

Does schema markup affect AI citations?

The rationale is strong — valid markup makes facts unambiguous to parse — but no major assistant has confirmed weighting it. It is good practice with a plausible mechanism, not a proven lever.

Can I submit my site to an AI assistant?

There is no submission process comparable to a sitemap. The closest equivalent is making sure you are crawlable, well-indexed and clearly written.

Why does an assistant describe my company incorrectly?

Usually because it is drawing on training data rather than live retrieval, or because your own site describes you ambiguously. Fix the second; the first resolves with model updates.

Sources

Every figure on this page traces to one of these. Dates are when we last read the page — pricing and features change, so treat anything older than a few months as a starting point rather than gospel. All outbound links here are nofollow.

  1. [1]
    Introduction to robots.txt

    Google Search Central · developers.google.com · Official documentation · read 2026-09-18

  2. [2]
    Rich Results Test

    Google · search.google.com · Vendor page · read 2026-09-18

Keep reading

This is the question everyone in the category wants answered and nobody can answer properly. The honest starting position is that the selection process is not documented, varies between products, cha…