llms.txt: controlling what AIs read on your site

For two years now, one small text file has absorbed an outsized share of GEO conversations: llms.txt. The promise is seductive: drop a machine-written map at the root of your site and take back control over what ChatGPT, Claude, Perplexity or Gemini read (and cite) about you. The 2026 reality is more nuanced: adoption on the site side has exploded, while support on the engine side remains marginal. In between sit a lot of marketing promises and very few server logs.

This article is a field report: what llms.txt really is, what AI crawlers actually do with it (numbers included), where real control lives (robots.txt and the block-or-feed trade-off) and why the file remains a reasonable bet, provided you take it for what it is.

A file that machines never request is not a control channel. It is an intention. Real control happens where crawlers actually go: robots.txt, content accessibility, authority.

llms.txt: the proposal

The idea comes from Jeremy Howard (Answer.AI), who formalized it in September 2024: a Markdown file at the site root (/llms.txt) offering language models a distilled version of your essential content. The logic is defensible: LLM context windows are limited, and your HTML pages are cluttered with navigation, scripts and cookie banners. Rather than letting the machine sort it out, you hand it a clean map.

The format is codified: an H1 title (the site name), a summary in a blockquote, then sections of annotated links, each link followed by one line saying what lives there. A variant, llms-full.txt, goes further by compiling the full content of key pages into a single file. Serious players have adopted the format (Anthropic, Stripe, Zapier, Cloudflare) and adoption grew nearly ninefold in a year. But almost always on the same site profile: technical documentation, built to be consumed by tools.

What AI crawlers actually do with it

This is where the field contradicts the marketing. Server-log analyses converge: roughly 97% of deployed llms.txt files receive no AI bot requests at all. One study covering more than 500 million AI crawler visits over 90 days found only a few hundred direct fetches of the file, a statistical drop in the ocean.

On the engine side, positions are public. Google does not support llms.txt and has no plans to. Gary Illyes confirmed it back in summer 2025, and John Mueller compared the file to the meta keywords tag, abandoned because self-declared signals cannot be verified. OpenAI and Anthropic officially point to robots.txt for managing crawler access, with no commitment to read llms.txt. Only Perplexity says it retrieves the file to help prioritize its reading. The format's real consumers today are elsewhere: coding agents (Cursor, Cline and friends) and integrations where a human explicitly feeds the file URL to a tool.

Interim conclusion, measured but firm: in 2026, publishing an llms.txt does not measurably improve your odds of being cited by ChatGPT, Gemini or Claude. Anyone selling you the opposite is selling superstition. Which does not make the file useless. More on that below.

The real control room: robots.txt versus AI crawlers

While llms.txt dominates the conversations, real control is exercised in a thirty-year-old file: robots.txt. Legitimate AI crawlers respect it, and that is where the only decision that matters gets made: who is allowed to read what. You just need to distinguish two families of bots with opposite roles: training bots, which collect content to feed future models, and answer bots, which index or fetch your pages to answer a user, with a citation and a link at stake.

User-agentWhat it doesIf you block it
GPTBotOpenAI model trainingYour content leaves future training sets; no direct effect on ChatGPT Search
OAI-SearchBotChatGPT Search indexYou disappear from ChatGPT Search results and citations
ChatGPT-UserLive fetches on a user's behalfChatGPT can no longer open your pages in session
ClaudeBotAnthropic collection (training)Same logic as GPTBot, on the Claude side
PerplexityBotPerplexity indexYou drop out of Perplexity's sourced answers
Google-ExtendedGemini (training and grounding)Affects neither your SEO nor AI Overviews
CCBotCommon Crawl (public corpora)You leave the corpora used by many models
BytespiderByteDance collectionLittle visibility upside; often blocked by default

The classic trap in that table: Google-Extended. Many site owners block it believing they are opting out of AI Overviews. They are not: AI Overviews and Google's AI mode rely on the classic Search index, fed by Googlebot. To disappear from them you would have to block Googlebot, meaning leaving Google altogether. Google-Extended only governs how Gemini uses your content.

Block or feed? The strategic trade-off

The right answer depends on your model, not on the trend. A publisher living off content licensing has a genuine protection calculation to make: its archive is its asset, and licensing deals negotiate better when access is not free. But for a services SMB, an e-commerce brand, a consultancy (the vast majority of my clients), the calculation flips: being cited by AIs is an emerging acquisition channel, and the visits it sends arrive pre-qualified by the answer that recommended you.

My field position, the one I apply in my GEO engagements: open the answer bots wide (OAI-SearchBot, PerplexityBot, ChatGPT-User), and decide deliberately on training bots depending on the nature of the content. A blog built to demonstrate expertise has every reason to feed the models, a proprietary knowledge base maybe not. What matters is that it is a decision, dated and documented, not a default setting inherited from a template.

Writing an llms.txt anyway (and writing it well)

So, should you publish one? Yes, the way you buy a lottery ticket at cost price. The bet is asymmetric: one to two hours of work, zero penalty risk, real consumers already today (coding agents, analysis tools, integrations), and a free option on a future where a major engine decides to read it. What is unreasonable is not the file; it is expecting a ranking effect from it.

The rules of a useful file: an H1 with the site name and one positioning line; a blockquote summary; three to five sections (offer, guides, case studies, contact) each holding a handful of links annotated with one line, not a sitemap in disguise; canonical URLs; an update at every significant publication. Keep llms-full.txt for sites whose documentation is the product. And put neither an exhaustive catalog nor keyword stuffing in it: the file speaks to machines that read well, not to an index you can manipulate.

What actually gets your site read (and cited) by AIs

If the lever is not the file, where is it? First, raw accessibility: most AI fetchers read HTML without executing JavaScript. Content that only exists after client-side rendering is invisible to them. Second, extractability: hierarchical headings, direct answers at the top of each section, lists and tables a model can quote without reconstruction, the principles I detail in my complete GEO guide. Then Schema.org structured data, still the best-supported machine-description channel. Finally, authority: generative engines cite the sources the web already cites, the bridge between link building and GEO that my SEO work covers.

And as always in GEO: measure before acting. Knowing whether ChatGPT, Perplexity or AI Overviews already cite you (and on which prompts) is a prerequisite to any crawl policy. That is exactly what my GEO audit and the monitoring I described in detail here are for.

The costliest mistake

Magical thinking, in both versions. The credulous version: publishing a plugin-generated llms.txt and considering the GEO job done, while the site remains unreadable without JavaScript and invisible in answers. The defensive version: blocking every AI bot as a protective reflex, then discovering six months later that competitors alone occupy ChatGPT's and Perplexity's answers on your business queries. In both cases the problem is the same: a default decision where an informed one was needed, bot by bot, content by content.

In the crucible: the map is not the territory

The lead, here, is the file published out of superstition, that no machine requests and nobody updates. The gold is a crawl policy chosen deliberately, content that fetchers read effortlessly, and authority that forces the citation. llms.txt can crown that edifice (a clean map, held out to whoever cares to take it), but it replaces none of its foundations.

If you want to know what AIs read (and above all what they cite) about your site today, that is precisely the first deliverable of my GEO expertise: a measured baseline, then a crawl and content policy that turns your pages into sources generative engines recommend.

Know what AIs cite about your site

I audit your presence in ChatGPT, Perplexity and AI Overviews, review your crawl policy (robots.txt, llms.txt) and deliver a prioritized GEO action plan.

Discover the GEO expertise → WhatsApp →

Related articles

GEO

Monitoring ChatGPT, Perplexity and AI Overviews in 2026

Resource
GEO

GEO: the complete 2026 guide

Resource
EXPERTISE

Service: GEO, getting cited by ChatGPT, Perplexity and Gemini

Resource