Back to Blog

The Data Debunk: Does llms.txt Actually Correlate With AI Citations?

Written by
Elsa JiElsa Ji
··9 min read
The Data Debunk: Does llms.txt Actually Correlate With AI Citations?

Roughly 10% of measured domains have shipped an llms.txt file since the format launched in September 2024. That adoption curve looks like a standard is forming.

But adoption and effect are two different questions. The one that actually matters for a GEO strategy is whether publishing llms.txt changes how often a brand gets cited in ChatGPT, Perplexity, or Google’s AI Overviews. Three independent studies, covering hundreds of thousands of domains, now have an answer.

What llms.txt Actually Is, Beyond the Hype

llms.txt is a Markdown file, hosted at a site’s root, that lists a brand’s most important pages in a clean, script-free format. Jeremy Howard and the team at Answer.AI proposed it on September 3, 2024, hosted at llmstxt.org. The pitch: a language model with a limited context window shouldn’t have to wade through navigation bars and ad scripts to find the content that matters.

The format itself is intentionally simple. An H1 title, an optional summary paragraph, then grouped links to key pages. Some sites also publish an expanded llms-full.txt that inlines the full text of every linked page into a single fetch.

That’s the whole spec. There’s no schema, no validator, no runtime API to register with. Anyone who can write a README can ship one in under an hour.

Why Everyone Assumed llms.txt Would Boost AI Citations

The marketing pitch around llms.txt borrowed credibility from two older files. robots.txt tells crawlers what not to touch. sitemap.xml tells search engines what exists. Both are widely respected because the crawler operators built compliance into their systems and said so publicly.

The Data Debunk: Does llms.txt Actually Correlate With AI Citations?

llms.txt skipped that step. No major AI platform has committed, on the record, to fetching it as part of how answers get generated.

That gap didn’t stop the analogy from spreading. Once a file looks like robots.txt for AI, it’s easy to assume it behaves like one too.

An assumption repeated enough times starts to look like a fact.

What the Data Actually Shows About llms.txt and Citations

This is the part the marketing pitch skips, and it’s where three separate research teams landed on the same conclusion.

SE Ranking’s analysis is the largest public study to date: roughly 300,000 domains, checked for llms.txt at the root, then measured against how often each domain got cited across major AI-powered answer engines. They ran two tests: a straightforward correlation analysis, then an XGBoost model trained with and without llms.txt as a feature. The model got slightly more accurate once llms.txt was removed, which in plain terms means the file was adding noise, not signal.

A second study from Trakkr scanned 37,894 AI-cited domains and cross-referenced 323,000-plus citations. The adoption rate came in at 12.7%, and the statistical test (Mann-Whitney U, chosen because citation counts are heavily skewed) returned a p-value of 0.81. That’s nowhere close to significant. Trakkr also found that among the top 50 most-cited domains, only 6% had adopted llms.txt at all, and adoption actually climbed further down the citation rankings. The sites hoping for a lift are shipping the file. The sites already winning citations mostly aren’t bothering.

Here’s the pattern across both datasets, side by side:

StudyDomains ScannedAdoption RateCitation Correlation
SE Ranking~300,00010.13%None found; removing the feature improved the prediction model
Trakkr37,89412.7%Not statistically significant (p=0.81)

Two different research teams, two different samples, two different statistical methods. Same null result.

Why LLMs May Be Ignoring llms.txt Entirely

The absence of a correlation stops looking surprising once you look at how these systems actually pull information.

ChatGPT’s web results run through Bing. Gemini runs through Google’s own index. Perplexity maintains its own crawl and index. None of the major consumer AI assistants have a separate, AI-specific discovery mechanism that visits a site’s root directory looking for a special file before generating an answer. They’re layered on top of search infrastructure that was built for a different purpose years earlier, then reads and synthesizes whatever that infrastructure surfaces.

That’s the mechanical reason llms.txt was always a longer shot than robots.txt. robots.txt works because it plugs directly into the crawler behavior it’s meant to influence. llms.txt asks a model to take a detour to a file most retrieval pipelines were never built to check.

Google’s own public position backs this up. At the Search Central Deep Dive event in Bangkok in July 2025, Gary Illyes stated plainly that Google does not use llms.txt and has no plans to. John Mueller drew a direct comparison to the keywords meta tag, a signal Google stopped trusting in the late 2000s because site owners control it and site owners can game it. In December 2025, an llms.txt briefly showed up on Google’s own developer documentation site and was pulled the same day. OpenAI, Anthropic, and Perplexity haven’t made an equivalent statement either way.

The Counterpoint Nobody Should Skip

Not every voice in this debate lands on the same side, and the strongest pushback is worth taking seriously rather than dismissing.

Wix’s AI Search Lab argues that Google’s own index contained between 30,000 and 60,000 llms.txt files as of October 2025, which they read as proof that Google is crawling the file even while saying otherwise. They also point out the format’s token efficiency: a clean Markdown index costs a fraction of the tokens a rendered HTML page does, which matters more as agentic workflows lean on tighter context budgets.

Both points are fair, and both come with a catch. A crawler visiting a file tells you it got fetched. It doesn’t tell you any model used that fetch to shape an answer. The same crawler indexes robots.txt and sitemap variants too, and indexing presence has never been the same thing as a ranking input.

The Data Debunk: Does llms.txt Actually Correlate With AI Citations?

Treat llms.txt as cheap insurance for a future where a major provider flips the switch, not as a lever that’s already paying off.

What Actually Correlates With AI Citations

If llms.txt isn’t the variable driving citations, the useful question becomes: what is?

The research points back to the fundamentals GEO practitioners already know. Content that’s structured for extraction, backed by clear entity signals, and already earning citations from other authoritative sources tends to show up more often in AI answers. None of that requires a special file. It requires the same substantive, well-organized content that’s always mattered, now read by a different kind of reader.

That’s less satisfying than a one-hour fix, but it’s what the data supports.

The practical problem is that most brands don’t actually know which of their pages are getting cited, or why. Guessing at causes wastes the same budget llms.txt already wasted for a lot of teams. Topify’s Source Analysis feature exists to close that gap. It tracks the exact domains and URLs that AI platforms cite when answering questions related to your category, so you can see which of your own pages are pulling weight and which competitor content is winning the citation instead.

How to Verify What’s Driving Your Own AI Visibility

Before spending another hour on llms.txt, it’s worth spending that same hour checking what’s already influencing your citation rate.

Start with a free GEO score check to get a baseline reading on how your site currently shows up across AI platforms. From there, Comprehensive GEO Analytics tracks visibility, sentiment, and position across ChatGPT, Gemini, and Perplexity over time, so a change in your content strategy shows up as a measurable shift rather than a guess. Pair that with Source Analysis to see which specific pages and domains AI systems are actually citing in your space right now.

That combination replaces speculation about file formats with a direct read on what’s working.

Conclusion

llms.txt is not a proven citation lever. Three studies covering hundreds of thousands of domains agree on that, and Google’s own public statements back it up. That doesn’t make the file harmful. It’s cheap to ship, and if a major provider ever does start using it, having an accurate one already in place costs nothing.

What it shouldn’t get is your GEO budget or your team’s attention as a primary strategy. The levers that correlate with AI citations today are the same ones that have mattered all along: structured, authoritative content that earns citations on its own merit. Spend the hour verifying what’s actually driving your visibility before spending it on a file the data says isn’t.

FAQ

Does llms.txt replace robots.txt or sitemap.xml? 

No. robots.txt and sitemap.xml are established standards that crawler operators have publicly committed to honoring. llms.txt is a proposal with no equivalent commitment from any major AI platform, so it doesn’t function as a replacement for either.

Do ChatGPT and Perplexity read llms.txt? 

There’s no public confirmation from OpenAI, Anthropic, or Perplexity that their retrieval systems fetch or weight llms.txt at runtime. Google has stated it does not use the file. Scattered third-party observations of bot traffic to the file exist, but none rise to an official commitment.

Is llms.txt worth setting up in 2026?

If it takes about an hour and you can keep it accurate, there’s little downside. Just don’t treat it as an AI visibility strategy or a paid line item, since the data available in 2026 shows no measurable citation lift from having one.

What actually influences whether AI platforms cite a brand? 

Structured, extractable content, clear entity signals, and existing citations from authoritative sources correlate with AI citation frequency far more than any single root-level file.

Read More

Topify dashboard

Get Your Brand AI's
First Choice Now