Guides · llms.txt

What is llms.txt, and does it matter?

7 min read · published 30 September 2026

llms.txt is a small Markdown file at the root of a site that tells language models what the site is and which pages to read first. It costs ten minutes to add and cannot hurt you. Whether it helps depends on who reads it, and the honest answer in 2026 is: fewer systems than the hype suggests, but more than none.

Where it came from

The proposal was published in September 2024 by Jeremy Howard of Answer.AI at llmstxt.org. The reasoning: a model that lands on a modern web page has to fight through navigation, scripts and boilerplate to find the content, and its context window is limited. A curated, plain-text index of the pages that matter lets it skip all that. The idea borrows the shape of robots.txt and sitemap.xml: one well-known file, one job.

The format

# Acme Store

> Handmade kitchenware, shipped across the EU.

Family-run since 2009. Prices include VAT.

## Docs
- [Shipping & returns](https://acme-store.com/shipping): costs, times, how to return
- [Size guide](https://acme-store.com/sizes): measurements for every product

## Optional
- [Company history](https://acme-store.com/about)

A companion file, llms-full.txt, holds the full text of the documentation in one file for systems that want everything at once. Serve both as text/plain or text/markdown. The llms.txt generator writes the file in this shape and validates an existing one.

Who reads it

This is where honesty is due. As of this writing, none of the large operators has documented that its crawler or assistant fetches llms.txt on its own. Google’s representatives have said their systems do not use it. OpenAI, Anthropic and Perplexity have published nothing that commits to it. So adding the file does not change how those systems crawl or rank you today.

It does get read in three places. Developer tools and coding assistants that let a user point them at a documentation site often look for llms.txt first, and many documentation platforms now generate one automatically. Agents given a URL and told to “read the docs” frequently fetch it because their authors know the convention. And people read it: it is a readable map of your site in a place developers know to look.

What it is not

llms.txt does not control anything. It does not allow or block crawlers; that is robots.txt. It does not keep pages out of AI answers; that is noindex, a Disallow for the search crawlers, or the legal signals described in the noindex guide. Treat it as a courtesy to readers, not a policy.

Why the checker gives it ten points

The checker awards 10 of 100 points for an llms.txt that exists and is served as text. That is not a claim that it moves rankings. It is a cheap, unambiguous signal that a site has thought about AI readers, and the 90 points that matter come from what crawlers can actually fetch.

Mistakes we see

Write one with the generator, then run the checker to confirm it is served correctly.

Check it on your site

More guides: How to block AI crawlers · robots.txt for AI crawlers · noindex vs robots.txt · Content-Signal explained · Cloudflare blocking Googlebot.

Spotted something out of date? Operators change their bots often; email info@canaicrawl.com and we’ll re-check the guide.