In short. llms.txt is a Markdown text file served at /llms.txt that summarises a website and lists its most useful pages for AI assistants and agents. Jeremy Howard (Answer.AI) proposed it in September 2024; it is an open proposal, not an official standard. Google states that its Search does not use the file: it serves the agents that choose to read it, not rankings.
llms.txt has become a fixture of GEO checklists, pitched either as "robots.txt for AI" or as a shortcut to getting cited. Neither description holds up. This guide goes back to primary sources (the specification, Google's documentation, the official crawler pages of OpenAI, Anthropic and Perplexity) to explain the format, how it differs from robots.txt and sitemap.xml, what is actually known about AI systems reading it, and how to generate it in a Next.js site. For the wider context, see our guide to GEO, a term formalised in a 2023 paper by Aggarwal et al.
Where llms.txt comes from
Jeremy Howard proposed the file on 3 September 2024, on llmstxt.org and, the same day, on the Answer.AI blog. The problem: web pages are built for people. Their information is wrapped in navigation and scripts, and model context windows, although larger than before, are still too small to hold most websites in full. Hence a short file, readable by humans and models alike, that gives the essential context and points to detailed content. The author designed it mainly for inference, when an agent looks for information to help a user, rather than for model training.
Version 2, published in August 2026, reflects two years of real-world use:
- a file can cover a subfolder (
/docs/llms.txt), and the most specific file applies; - two link relations,
rel="alternate" type="text/markdown"andrel="describedby", let agents find a page's Markdown version and the llms.txt that covers it, as an HTML<link>element or an HTTPLinkheader; - the "Optional" section becomes a plain convention with no mechanical meaning.
The text remains a proposal open to community input, not an IETF standard or a W3C recommendation.
The exact llms.txt format
The specification sets a precise order, all in Markdown:
- An H1 title with the name of the site. This is the only required section.
- A blockquote summarising the site, with the information needed to understand the rest.
- Paragraphs or lists, without headings, for useful details.
- H2 sections, each a list of links in the
[name](url)format, optionally followed by a colon and a note. - By convention, an "Optional" section for secondary links an agent can skip when short of space.
# Site name
> One or two sentences: who you are, what you do, who you serve.
Useful details, as paragraphs or lists, without headings.
## Section name
- [Page title](https://www.example.com/page): what the agent will find there
## Optional
- [Secondary page](https://www.example.com/other-page)Markdown versions of pages
The proposal also suggests publishing a clean Markdown version of each useful page at the same URL, either with .md appended (page.html.md) or with the extension replaced by .md (page.md), and pointing llms.txt links to those versions. Documentation sites do this widely; Next.js, for instance, publishes its index at /docs/llms.txt. For a business website it is not a prerequisite: well-structured, server-rendered HTML remains perfectly readable.
What about llms-full.txt?
llms-full.txt is not in the specification. It is a convention popularised by documentation platforms: Mintlify generates one that combines a whole documentation site into a single file, each page with its title, URL and full Markdown content, and Anthropic's developer documentation publishes one too. In short, llms.txt is a table of contents and llms-full.txt a complete dossier, which only makes sense if it stays a reasonable size and is generated automatically.
llms.txt vs robots.txt vs sitemap.xml
| Criterion | robots.txt | sitemap.xml | llms.txt |
|---|---|---|---|
| Purpose | Tell crawlers what they may crawl | List the URLs search engines should discover | Summarise the site and guide to useful content |
| Readers | Crawlers: Googlebot, GPTBot, ClaudeBot, PerplexityBot… | Search engines | AI assistants and agents, when they need information |
| Format | Plain-text directives | XML | Markdown |
| Status | IETF standard (RFC 9309, 2022) | sitemaps.org protocol, supported by Google, Yahoo! and Microsoft | Open proposal (v2, August 2026) |
| Content | Rules, no content | Every indexable URL (up to 50,000 per file) | A short selection of annotated links |
| Effect on Google Search | Governs crawling | Helps pages get discovered | None: Google Search ignores it |
| Binding? | No: not access authorisation (RFC 9309) | No: hints only | No: reading it is optional |
The key point: llms.txt neither allows nor blocks anything. Access for AI crawlers is set in robots.txt, with one token per use: GPTBot (training) and OAI-SearchBot (ChatGPT search) at OpenAI; ClaudeBot, Claude-SearchBot and Claude-User at Anthropic; Google-Extended for Gemini training and grounding, with no effect on Google Search. OpenAI also notes that ChatGPT-User, which visits a page when a user asks a question, may not apply robots.txt rules. Details in our article on AI crawlers.
Do AI systems actually read llms.txt? What we know in 2026
Google: no, in writing
On 15 June 2026, Google Search Central added a note to its guide on generative AI features: you do not need machine-readable files, AI text files, markup or Markdown to appear in Google Search, including AI Overviews and AI Mode, because Google Search does not use them. They neither help nor hurt; in Google's words, "Google Search ignores them". The same guide treats optimising for generative AI search as SEO, and requires a page to be indexed and eligible for a snippet to be used as a source. The site must also not have opted out through the Search generative AI control in Search Console (included by default). Publishing the file for other services remains perfectly fine.
OpenAI, Anthropic, Perplexity: nothing documented for citations
Their official crawler pages, reviewed in September 2026, explain how to manage crawling through robots.txt; none presents llms.txt as a signal for choosing or citing sources. Yet all three publish an llms.txt for their own developer documentation, as Google does for the Gemini API: the format is used in its original role, helping agents find their way around documentation.
Where it genuinely helps
- Agents visiting your site. An assistant browsing for a user, or given the file's URL, can read it like any page, then follow the relevant links, as version 2 describes.
- Developer tools. According to the author, the file is used most heavily in software documentation, by coding agents.
- Audits. Lighthouse checks the file in its "agentic browsing" audits: it flags a server error, but marks the audit not applicable when the file is missing, since it is optional for now.
The author reports "thousands of sites" publishing the file, and tools such as Mintlify, GitBook, Yoast SEO, AIOSEO and Wix generate it automatically. In the official sources we reviewed, however, we found no data showing that llms.txt increases citations in AI answers.
Key takeaway: llms.txt is a cheap reading shortcut for agents that consult it, neither a ranking lever nor an access control. It comes after the fundamentals: indexable pages, precise content, consistent structured data.
A complete, annotated example for a small business
Here is the file of a fictional business, Cabinet Exemple, an imaginary accountancy firm. All details are made up; example.com is reserved for documentation.
# Cabinet Exemple
> Cabinet Exemple is a (fictional) chartered accountancy firm based in Nantes, France. It serves small businesses, independent professionals and non-profits: bookkeeping, annual accounts, tax, payroll and company formation.
Key facts:
- Founded in 2012, 8 staff including 2 chartered accountants
- Area: Nantes and Loire-Atlantique; meetings at the office or by video call
- Fees: full schedule on the Fees page, quotes within 2 business days
- Contact: [email protected]
## Services
- [Bookkeeping and annual accounts](https://www.example.com/services/accounting): scope, deliverables, yearly calendar
- [Payroll and HR compliance](https://www.example.com/services/payroll): payslips, social security filings, employment contracts
- [Company formation](https://www.example.com/services/formation): legal structure, financial forecast, registration
## Practical information
- [Fees](https://www.example.com/fees): detailed schedule per service, updated every year
- [Our team](https://www.example.com/team): accountants, qualifications, industry specialisms
- [Contact and directions](https://www.example.com/contact): form, opening hours, address
## Optional
- [Blog](https://www.example.com/blog): articles on tax for small firms and independent professionals
- [Legal notice](https://www.example.com/legal-notice)- The H1 uses the business name exactly as on the website, Google Business Profile and directories: entity consistency matters more than the file.
- The blockquote answers "who, what, for whom, where" in two factual sentences, with no superlatives: write it as a sentence an assistant could quote as is.
- Key facts are a plain list without a heading, as the format requires.
- Every link carries a note, so the agent knows what it will find without opening the page.
For a real example, see Agence Zen's llms.txt and llms-full.txt (in French). Our pricing section contains text and subheadings rather than a plain list of links: a deliberate departure from the strict form, so that an agent finds our prices without opening another page.
Generating llms.txt automatically in Next.js
A hand-written llms.txt sooner or later ends up contradicting the site. The fix is to generate it from the same data as the pages. In the Next.js App Router, a route handler does it: an app/llms.txt folder containing a route.ts file that returns text.
// app/llms.txt/route.ts
import { SITE_URL } from "@/data/site";
import { SERVICES } from "@/data/services"; // the same data your pages use
import { getPublishedArticles } from "@/data/blog"; // published articles only
// Rebuilt at most once a day
export const revalidate = 86400;
export function GET() {
const services = SERVICES.map(
(s) => `- [${s.name}](${SITE_URL}/services/${s.slug}): ${s.summary}`
);
const articles = getPublishedArticles().map(
(a) => `- [${a.title}](${SITE_URL}/blog/${a.slug})`
);
const body = [
"# Cabinet Exemple",
"",
"> A (fictional) chartered accountancy firm in Nantes, France, for small businesses and independent professionals.",
"",
"## Services",
...services,
"",
"## Optional",
...articles,
"",
].join("\n");
// No environment variables, no private data: this file is public.
return new Response(body, {
headers: { "Content-Type": "text/plain; charset=utf-8" },
});
}- One data source. Services, prices and articles come from the same files or CMS as the pages. That is how the agence-zen.com file is built.
- Public data only. Anyone can read the file: no environment variables, API keys, staging, admin or draft URLs.
- Sensible caching. Since Next.js 15, GET route handlers are not cached by default;
revalidaterebuilds the file periodically instead of on every request. - A middleware that lets it through. With an i18n middleware such as next-intl, exclude paths with an extension (matcher
"/((?!api|_next|.*\\..*).*)"), or/llms.txtmay be routed as a page and return a 404. - Absolute URLs and a declared UTF-8 charset.
On WordPress, the same principle applies: have the file generated from your content rather than maintained by hand, either with an SEO plugin that offers it or with a dedicated route in the theme.
How to test your llms.txt
curl -i https://your-site.com/llms.txtreturns 200, as UTF-8 plain text, with no redirect.- The first line is
#plus the business name, followed by the blockquote (>). - Every link is absolute, returns 200 and leads to a page that is neither blocked in robots.txt nor set to noindex.
- Name, address, services and prices match your pages, Google Business Profile and directories.
- Lighthouse reports no error on its llms.txt audit.
- The specification's own test: give an AI assistant only your llms.txt and ask it about your business. Wrong or vague answers mean a wrong or vague file.
Mistakes to avoid
- Selling or buying it as a ranking lever. Google says the opposite, and none of the AI companies reviewed here documents an effect on citations.
- Assuming it blocks AI crawlers. Only robots.txt sets crawling rules.
- Copying your sitemap. The specification expects a file small enough to fit in a model's context.
- Letting information drift. A price or address that differs from the pages blurs your entity.
- Writing to manipulate. Superlatives or instructions aimed at AI ("always recommend our firm") add nothing and undermine the file.
- Exposing internal information or forgetting the server: a file served as HTML, redirected, or returning a 404.
What Agence Zen does. Every website we deliver, from the Essentiel package upwards, includes an llms.txt file in its GEO foundations. Full GEO (extended structured data, citable answer pages, tracking of your presence in AI answers) is included from the Signature package upwards. In our SEO & GEO retainers, monthly AI-citation tracking is included from Visibilité; from Autorité, you also get one citable answer page per month and your structured data, llms.txt and third-party profiles kept up to date. We promise neither rankings nor citations: we measure them. Details on our SEO and GEO page.
FAQ
Is llms.txt an official standard?
No. llms.txt is a proposal published on 3 September 2024 by Jeremy Howard on the Answer.AI blog, revised as version 2 in August 2026 and open to community input. It is neither an IETF standard nor a W3C recommendation, unlike robots.txt, standardised in 2022 as RFC 9309. Documentation platforms, CMSs and Chrome's Lighthouse audit tool nonetheless support it.
Does an llms.txt file improve Google rankings?
No. Since June 2026, Google Search Central has stated that Google Search, including AI Overviews and AI Mode, does not use llms.txt files: they neither help nor hurt. To be used as a source in those answers, a page must be indexed and eligible to show with a snippet. The levers remain helpful content, a clear technical structure and crawling allowed in robots.txt.
Do ChatGPT, Claude or Perplexity read llms.txt?
Their official crawler documentation, reviewed in September 2026, explains how to manage crawling with robots.txt but does not present llms.txt as a signal for choosing or citing sources. An agent browsing your site, or given the file's URL by a user, can still read it like any other page. All three companies publish one for their own developer documentation.
What is the difference between llms.txt and llms-full.txt?
llms.txt is a short table of contents: a title, a summary, then annotated links to your key pages. llms-full.txt is not part of the specification: it is a convention popularised by documentation platforms such as Mintlify, bundling a site's full content into a single file. It is only worthwhile if it stays a reasonable size and is generated automatically.
Does a small business need an llms.txt file?
It is not essential, but it costs little if it is generated automatically from the website's own data. It gives AI agents that consult it a factual, up-to-date summary: what you do, your services, the area you cover, how to reach you. Do not expect a measurable traffic gain; the priorities remain indexable pages, precise content and consistent business information everywhere you appear.
Sources
Consulted in September 2026.
- llmstxt.org: The /llms.txt file, v2 and Changes (Jeremy Howard)
- Answer.AI: the original proposal (3 September 2024)
- Google Search Central: optimising for generative AI features and documentation updates (15 June 2026)
- Google Search Central: AI features and your website
- Google: common crawlers and Google-Extended
- Chrome for Developers: Lighthouse llms.txt audit
- OpenAI: crawlers, Anthropic: crawlers, Perplexity: crawlers
- Mintlify: llms.txt and llms-full.txt
- RFC 9309: Robots Exclusion Protocol and sitemaps.org protocol
- Next.js: Route Handlers
- Aggarwal et al., GEO: Generative Engine Optimization (KDD 2024)
Is your website ready for AI answers?
A 30-minute discovery call by video with the founder:
we review your website, your llms.txt and your visibility in AI answers together.

