Guides · llms.txt
What is llms.txt, and does it matter?
llms.txt is a small Markdown file at the root of a site that tells language models what the site is and which pages to read first. It costs ten minutes to add and cannot hurt you. Whether it helps depends on who reads it, and the honest answer in 2026 is: fewer systems than the hype suggests, but more than none.
Where it came from
The proposal was published in September 2024 by Jeremy Howard of Answer.AI at llmstxt.org. The reasoning: a model that lands on a modern web page has to fight through navigation, scripts and boilerplate to find the content, and its context window is limited. A curated, plain-text index of the pages that matter lets it skip all that. The idea borrows the shape of robots.txt and sitemap.xml: one well-known file, one job.
The format
# Acme Store
> Handmade kitchenware, shipped across the EU.
Family-run since 2009. Prices include VAT.
## Docs
- [Shipping & returns](https://acme-store.com/shipping): costs, times, how to return
- [Size guide](https://acme-store.com/sizes): measurements for every product
## Optional
- [Company history](https://acme-store.com/about)
- An H1 with the site or project name. This is the only required line.
- A blockquote with a one-line summary.
- Any number of plain paragraphs with context.
- H2 sections, each a list of
- [title](url): noteslines. Links should be absolute, and ideally point at pages that are readable as text or Markdown. - A section named Optional has a special meaning: a reader short on context may skip it.
A companion file, llms-full.txt, holds the full text of the documentation in one file for systems that want everything at once. Serve both as text/plain or text/markdown. The llms.txt generator writes the file in this shape and validates an existing one.
Who reads it
This is where honesty is due. As of this writing, none of the large operators has documented that its crawler or assistant fetches llms.txt on its own. Google’s representatives have said their systems do not use it. OpenAI, Anthropic and Perplexity have published nothing that commits to it. So adding the file does not change how those systems crawl or rank you today.
It does get read in three places. Developer tools and coding assistants that let a user point them at a documentation site often look for llms.txt first, and many documentation platforms now generate one automatically. Agents given a URL and told to “read the docs” frequently fetch it because their authors know the convention. And people read it: it is a readable map of your site in a place developers know to look.
What it is not
llms.txt does not control anything. It does not allow or block crawlers; that is robots.txt. It does not keep pages out of AI answers; that is noindex, a Disallow for the search crawlers, or the legal signals described in the noindex guide. Treat it as a courtesy to readers, not a policy.
Why the checker gives it ten points
The checker awards 10 of 100 points for an llms.txt that exists and is served as text. That is not a claim that it moves rankings. It is a cheap, unambiguous signal that a site has thought about AI readers, and the 90 points that matter come from what crawlers can actually fetch.
Mistakes we see
- A 200 HTML page. Single-page apps answer
/llms.txtwith the home page. The checker reports “200, but an HTML page” and does not count it. - Relative links such as
/docs, which a model reading the file out of context cannot resolve. - The whole sitemap. Hundreds of links defeat the purpose. Ten to forty pages that answer real questions is the sweet spot; put the rest in llms-full.txt or under Optional.
- No summary. The blockquote is the line assistants quote when describing the site. Write it.
Write one with the generator, then run the checker to confirm it is served correctly.
Check it on your site
More guides: How to block AI crawlers · robots.txt for AI crawlers · noindex vs robots.txt · Content-Signal explained · Cloudflare blocking Googlebot.
Spotted something out of date? Operators change their bots often; email info@canaicrawl.com and we’ll re-check the guide.