llms.txt Example: Found but 2 Broken Links and Stale Content
This example shows the aiwebpageseo llms.txt Audit for a SaaS marketing site. llms.txt exists (HTTP 200) with a readability score of 86 of 100. 14 link targets, 2 returning 404, and 3 references contain outdated information. llms-full.txt missing — recommended improvement that adds ~12% AI citation rate on average.
/docs/api.md returns 404
Linked from llms.txt but API docs moved to /developers/api/ in March. Update llms.txt link OR add a 301 redirect at /docs/api.md → /developers/api/.
/policies/ai.md returns 404
Policy URL incorrect — actual path is /policies/ai-usage.md. Update llms.txt so the link resolves. No claim is made here about citation likelihood: nobody outside the assistant operators knows what affects it, and a linked policy page is not a documented factor. The reason to fix it is simpler — a file whose links 404 is a file that has not been maintained, and it is advertising that fact to every client that reads it.
"47 supported tools" outdated
Integrations page now lists 62 tools. The file states 47. There is no "AI confidence" score being computed anywhere — the concrete problem is that a machine-readable file on your own domain now contains a false statement about your product, and any client that reads it is being misinformed by you.
No llms-full.txt
Optional. It provides expanded content for each linked document in one file, which spares a client from fetching each page separately. There is no evidence that publishing it raises citation rates, and any figure claiming otherwise — including one derived from a vendor cohort — is a correlation with no control for the far likelier explanation, which is that sites bothering to publish one are sites that maintain their content generally.
| Change | Effort | Impact |
|---|---|---|
| Fix /docs/api.md link → /developers/api/ | Trivial | High |
| Fix /policies/ai.md link → /policies/ai-usage.md | Trivial | High |
| Update integration count to current value | Trivial | Medium |
| Add llms-full.txt with expanded summaries | Medium | High |
| Add "Updated:" date at top of file | Trivial | Medium |
| Add explicit AI training policy URL | Low | Medium |
Before reading a score of 86 out of 100 as an achievement, it is worth being clear about what is being scored. llms.txt is a community proposal, not a standard. No assistant operator publishes a specification for it, none has stated that it reads one, and adoption across the industry is partial and unverified. Its effect on whether you get cited has never been demonstrated by anybody.
That means every claim you will read about llms.txt raising citation rates is unsupported. Including, to be plain, claims made in reports like this one: a cohort figure showing that sites with an llms-full.txt get cited more often has no control for the obvious confounder, which is that a team disciplined enough to publish one is a team that maintains its content generally. The file did not cause the outcome; the same habit produced both.
What the file is: a short, curated, human-written map of your site, in Markdown, at a predictable path. It costs almost nothing to publish and it does no harm. Publish it on that basis and you will not be disappointed by it.
A curated map exists to be reliable. Two of fourteen links lead nowhere, which means roughly one in seven of the destinations this file recommends does not exist. There is no scoring system docking you points for that. The cost is more direct: any client that follows those links has wasted a fetch and learned that the file is stale, and the file's only purpose was to be trusted.
Both failures also point at something larger. /docs/api.md returns 404 because the API documentation moved in March and nobody updated the things pointing at it — which means llms.txt is almost certainly not the only reference that broke. Fix the file, and then put a 301 at the old path, because whatever else links there is equally broken and you have not looked.
The stale figure has the same shape. The file says 47 supported tools; the integrations page says 62. No confidence score is being computed anywhere. The concrete problem is that a machine-readable file on your own domain now states something false about your own product, and it will keep stating it until somebody notices.
These two files get conflated constantly, and the confusion produces sites that believe they have controlled something they have not.
Governs access, by convention
It tells named crawlers what they may fetch. It is honoured by operators who choose to honour it, which is most of the reputable ones, and it is the correct place to name GPTBot, Google-Extended, ClaudeBot, OAI-SearchBot and the rest. It is a request, not an access control — if you need enforcement, that is a WAF rule or an authentication wall, and it will apply to everyone.
Recommends content, and controls nothing
It says "these are the pages worth reading". It makes no access claim, prevents no fetch, and a permissions statement written inside it has no mechanism behind it and no established legal weight. Writing "no AI training" into llms.txt does not stop AI training.
The consequence matters. If you want to decline model training, that is a robots.txt decision naming GPTBot and Google-Extended — and it is a separate decision from retrieval. Blocking GPTBot does not remove you from ChatGPT's cited answers, because those are served by different agents. Disallowing Google-Extended has no effect on Google Search or on AI Overviews, which run on ordinary Googlebot access.
And whatever you write in either file, verify in your access logs that reality matches. A permissive robots.txt sitting behind a WAF rule that drops anything with "bot" in the user-agent produces a perfect audit and zero crawler traffic.
The temptation with a machine-readable file is to be exhaustive. Resist it: an llms.txt listing two hundred URLs is a sitemap with prose attached, and it has thrown away the only thing that made it different, which was curation.
The right question is the one a new colleague would ask: if somebody has time to read six pages of this site and no more, which six, and why? Answer that in plain sentences. A site name, a description of what the site is actually for, and a short list of links each with one honest sentence explaining what it contains and who would want it.
The links should point at the pages that explain things, not the pages that sell things. A retrieval system summarising your documentation will find the documentation useful; it will find your homepage hero copy useless, because marketing prose does not survive being lifted out of its layout. Point at the substance.
Then give the file an owner. Nothing breaks visibly when llms.txt rots — no monitor fires, no user complains, the site keeps working — which is precisely why this one is now advertising a 404 and a number that is fifteen tools out of date. Regenerate it whenever documentation ships, and add its links to whatever already checks the site for broken URLs.
If the goal behind publishing llms.txt is to be readable by AI systems, then it is worth knowing where that goal is genuinely won and lost — because it is not in this file.
Can the retrieval agents actually fetch you?
Confirmed in your access logs, not in your robots.txt. A silent WAF block is the single most common cause of an assistant never seeing a site, and it produces no error anybody notices.
Is the substance in the initial response?
Most retrieval systems parse rather than render. Content assembled client-side after a fetch resolves is, to them, content that does not exist — including your prices, your feature lists and your FAQ answers.
Does your markup say what your text means?
Accurate JSON-LD identifies a price as a price and an author as an author. It clarifies; it does not persuade, and there is no verified evidence that denser schema wins citations.
Have you written the answer?
The only item with real headroom. Assistants surface passages that stand up when lifted out of the page. Everything else on this list is a precondition for that passage being found; none of them is a substitute for writing it.
Is llms.txt an official standard that AI companies follow?
No. It is a community proposal: a Markdown file at your site root listing your most useful pages with short descriptions, so a client with a limited context window can be pointed at the good material rather than crawling everything. It is not an IETF standard, no assistant operator publishes a specification for it, and adoption is partial and unverified. Its effect on citation has never been demonstrated. That is not an argument against publishing one — it costs almost nothing and it forces a useful editorial decision about which of your pages actually matter — but it is an argument against expecting anything from it.
Do the two 404s in the file actually matter?
Yes, though not for the reason usually given. A client that follows a link in your llms.txt and receives a 404 has learned that the file is out of date, and it has wasted a fetch it might not repeat. There is no scoring mechanism being docked. The real cost is that the file's entire purpose is to be a reliable map of your site, and a map with two roads leading nowhere is not one. Fix the paths, or put 301 redirects at the old ones — which you should do anyway, since anything else linking to /docs/api.md is equally broken.
Is llms.txt the same thing as robots.txt?
No, and conflating them causes real mistakes. robots.txt is a long-standing convention that tells crawlers what they may and may not fetch; it is honoured by the operators who choose to honour it, and it governs access. llms.txt makes no access claim at all — it is a curated index, a recommendation about what is worth reading. Blocking a crawler is done in robots.txt or at your WAF. Nothing in llms.txt prevents anything from being fetched, and a permissions statement inside it has no mechanism behind it.
Does llms.txt control whether my content is used for AI training?
It does not. Training access is governed by the crawler directives in robots.txt — GPTBot for OpenAI training, Google-Extended for Gemini training — and, if you need enforcement rather than a polite request, by your WAF. A statement inside llms.txt saying that your content may not be used for training has no technical effect and no established legal weight. If declining training use matters to you, name the agents in robots.txt, and remember that these are separate decisions from retrieval: blocking GPTBot does not remove you from ChatGPT's cited answers, which are served by different agents.
What should actually go in the file?
The pages you would want somebody to read if they had time for six of them and no more. That is the whole discipline, and it is why the file is worth writing even if no machine ever reads it. A site name, a description of what the site is for, and then a curated list of links with a sentence each explaining what each one contains and who it is for. Keep it short. A file listing two hundred URLs is a sitemap with prose attached, and it defeats the point — which was to say what matters, not to enumerate what exists.
How often does it need updating?
Whenever the pages it points at change, which in practice means it needs an owner. This file is unusual in that nothing breaks visibly when it rots — the site keeps working, no monitoring fires, no user complains — and so it quietly drifts out of date, exactly as this one has. The pragmatic answer is to regenerate it as part of whatever process publishes your documentation, and to add its links to whatever already checks your site for broken URLs. A file nobody owns will be wrong within a year.
If the benefit is unproven, why publish one at all?
Because the cost is a few minutes and the downside is nil, and because the exercise itself is worth more than the artefact. Deciding which eight pages represent your site, and writing one honest sentence about each, is a question most teams have never sat down and answered. The file that falls out of it is a byproduct. If assistant operators do converge on reading it, you have it already; if they do not, you have spent an afternoon getting clear about what your site is for, which was never wasted time.
Related Demo Reports
Run llms.txt Builder on Your Own Site
Get your real llms.txt audit with link health checks, readability scoring and ready-to-paste improvements.