Jack F, Co-founder
Part of the Answyn AEO Handbook.
Now readingHow to grade AEO advice
Brands earn citations in AI answers by being reachable, relevant to the exact question, and findable in the places engines already trust. Ranking still helps, but it is not the gate: only 38% of AI Overview citations in Ahrefs' 2026 analysis came from pages in the top 10. Schema and llms.txt do not move the number. Check first that you are not blocking crawlers or snippets.
Search "how to get cited by ChatGPT" and you will get a hundred listicles telling you roughly the same twelve things. Some of those things are confirmed in writing by the companies that build the engines. Some come from studies with real sample sizes. Some are one person's experience repeated until it sounded like a law. And a few have been tested properly and found to do nothing at all.
The problem is not that the advice is wrong. It is that it all arrives at the same volume, so you cannot tell which bits to spend money on.
This piece sorts it. Every claim below carries a grade for how well evidenced it is, and we have included the ones that failed, because knowing what not to buy is worth as much as knowing what to build.
How should you grade AEO advice?
On four levels, and you can apply them yourself to anything you read. [A] the engine says so in writing. [B] a published study with a stated sample found it. [C] practitioners keep observing it but nobody has tested it properly. [D] everyone repeats it and no source exists. We call these the AEO Evidence Grades, and we use them throughout the handbook.
| Grade | What it means | How much to trust it |
|---|---|---|
| [A] Platform-confirmed | Documented by Google, Microsoft, OpenAI or Anthropic, or said on the record by someone who works on the product | Act on it |
| [B] Study-supported | A published study with a stated sample and method you can go and read | Act on it, but check the sample and who paid for it |
| [C] Practitioner evidence | Repeatedly observed by people doing the work, no controlled test | Worth trying, don't build a strategy on it |
| [D] Untraceable | Widely repeated, no source that resolves to a number | Ignore until someone measures it |
Two things to watch even at grade B. Most AEO research is published by companies selling AEO tools, ourselves included, and almost none of it has a control group. A study that only looks at pages that got cited cannot tell you what makes a page get cited, because it never looked at one that did not.
What does Google actually say about optimising for AI answers?
That there is nothing new to do. Google published a formal guide in 2026 and its position is that generative AI features run on the same core ranking systems as ordinary Search, so ordinary SEO is the optimisation. [A]
Worth quoting directly, because a lot of money is currently being spent on things this rules out (Google Search Central):
- "You don't need to create new machine readable files, AI text files, markup, or Markdown."
- "Structured data isn't required for generative AI search."
- "There's no requirement to break your content into tiny pieces for AI."
- "You don't need to write in a specific way just for generative AI search."
What they do say helps: unique, non-commodity content with a point of view rather than a rearrangement of what is already out there; pages that are crawlable and indexed; and, specifically, pages eligible to be shown with a snippet. That last one is the single most actionable sentence Google has published on this, and we will come back to it.
They also warn against two things directly: spinning up a separate page for every possible query variation, which trips their spam policies, and chasing inauthentic mentions across the web.
One honest caveat. This is Google describing Google. It covers AI Overviews and AI Mode. ChatGPT, Claude, Perplexity and Copilot publish nothing equivalent, and there is no reason to assume they weight things identically.
Do you need to rank to be cited?
Less than you used to, and the change has been fast. In Ahrefs' March 2026 analysis of 863,000 keywords and 4 million AI Overview URLs, 38% of cited pages appeared in Google's top 10 organic results. In their July 2025 version, it was 76%. [B]
The rest splits almost evenly: 31.2% of cited pages ranked between positions 11 and 100, and 31.0% ranked beyond position 100 or not at all (Ahrefs, via Search Engine Journal).
Be careful with the 76 to 38 story, though. Ahrefs themselves say part of the drop comes from better parsing in the newer study, and that the two datasets are not directly comparable. So treat it as "top-10 ranking is not the gate people assume" rather than "citation halved its dependence on rank in eight months".
Either way, the practical read holds. Nearly two thirds of AI Overview citations go to pages that are not in the top 10 for the query being asked. If your AEO plan is "rank first, get cited second", you are queueing for a door that is not the main entrance.
Which factors correlate with being cited?
The most thorough synthesis available reviewed 54 published studies, experiments and patents and scored 23 factors on how repeatable and well-evidenced each one is. The top of the list is unglamorous: can the engine reach your URL, do you rank, and does your page answer the exact question. [B]
Cyrus Shepard's analysis scored each factor from 2 to 9.5 on three criteria: how often the same finding recurs across studies, how large the underlying datasets are, and whether any official documentation or patent supports it (Zyppy).
The top five:
| Factor | Score |
|---|---|
| URL accessibility | 9.5 |
| Search rank | 9.4 |
| Fan-out rank | 9.3 |
| Preview control | 9.2 |
| Query-answer match | 9.2 |
And the bottom of the list: llms.txt, at 2.0, with the note that "we're unable to find any credible evidence or experiments".
Shepard's own conclusion is the useful one: winning SEO usually means winning AI citations, with extra steps. His caveat is equally important and we will repeat it, because it applies to nearly everything in this field: these are not ranking factors in the traditional sense, and correlation is not causation.
Notice what the top five have in common. Four of them are about being reachable and being relevant. None of them is a tactic you buy.
What content structure actually gets cited?
Structure moves the number on its own, independently of what the page says. A controlled experiment across six generative engines found that structural optimisation alone lifted citation rate by 17.3%, while holding semantic quality constant. [B]
The paper decomposes structure into three levels: macro (how the document is laid out), meso (how information is organised within it) and micro (visual formatting). Optimising across all three improved both citation rate and rated quality, by 17.3% and 18.5% respectively (arXiv 2603.29979).
The grade-B caveat is real: the published abstract does not specify which structural elements were manipulated, so you cannot reproduce it exactly from the paper.
What practitioners consistently find alongside it, at grade C: an answer that stands alone in the first two or three sentences under each heading, headings phrased as the question a person would actually ask, and tables for anything comparative. The logic is that engines do not lift a page, they lift a passage, so a passage that only makes sense with the four paragraphs above it is a passage that cannot be lifted.
There is a useful sanity check here that costs nothing. Read one section of your page on its own, out of context. If it does not answer a question by itself, an engine cannot use it by itself either. Our answer readiness checker does this pass automatically if you would rather not eyeball it.
Does schema markup help AI citations?
No, on the best evidence available, and this is the clearest null result in the field. [B]
Ahrefs tracked 1,885 pages that added JSON-LD schema between August 2025 and March 2026 and compared them against 4,000 control pages matched on prior citation levels, using a difference-in-differences design (Ahrefs):
| Platform | Change after adding schema |
|---|---|
| Google AI Overviews | −4.6% (small, statistically significant, wrong direction) |
| Google AI Mode | +2.4% (indistinguishable from zero) |
| ChatGPT | +2.2% (indistinguishable from zero) |
The correlation everyone quotes comes from the same study: 53% of AI-cited pages carry JSON-LD, about three times the rate of uncited pages. That is confounded. Sites that implement schema also write better content, earn more mentions and maintain their pages, and the difference-in-differences design exists precisely to strip that out. When you strip it out, the effect goes.
In a companion test, none of five major AI systems parsed JSON-LD at all when fetching pages live.
Caveats worth stating: the sample was pages already being cited, so schema might still help a page that has never been picked up; the measurement window was 30 days; and all schema types were pooled, so it is possible some help and others hurt.
Schema is still worth having for rich results, knowledge graph entity resolution and voice. It just is not an AI citation lever, and Google says as much in the guide quoted above.
Does llms.txt do anything?
No. It is the one thing in this piece that is both widely recommended and explicitly ruled out by the platform. [A]
Google's own words: llms.txt files "will neither harm nor help your site's visibility". Shepard's synthesis scored it 2.0 out of 9.5, the lowest of 23 factors, on the grounds that no credible experiment supports it.
We still ship one at answyn.com/llms.txt, and we would rather be straight about why: it takes twenty minutes, it does no harm, and if any engine ever starts reading it we would like to already be there. That is a cheap option, not a strategy. If someone is charging you for llms.txt implementation as a line item, that is the tell.
How fresh does your content need to be?
Fresher than average, but by much less than the headline statistics suggest, and the thing that matters is the update date rather than the publish date. [B]
The most quoted figure in AEO is that 88% of AI-cited pages are from the last two years. That figure measures when pages were last updated, not when they were published, and in the same dataset only 42% had been published within the past year (Seer Interactive).
With a control group, the effect shrinks considerably. Across 16.975 million cited URLs, AI assistants cite content averaging 2.9 years old against 3.9 years for organic results, and Google AI Overviews actually cites slightly older content than organic (Ahrefs).
And the counterintuitive bit: pages cited consistently month after month are less fresh than pages that appear once and vanish. Recency buys a spike. Maintenance buys a habit.
We have written this one up properly, including what each engine actually reads when it works out your last-modified date, in content freshness and AI citations.
Where do AI citations actually come from?
Substantially from places you do not own. In a 30 million source analysis by Peec AI across ChatGPT, Google AI Mode, Gemini, Perplexity and AI Overviews, the most cited domains were Reddit, YouTube, LinkedIn, Wikipedia and Forbes. [B]
Per-engine, the same study found ChatGPT leaning on Wikipedia, Reddit and editorial sites, Google leaning on platforms including Facebook and Yelp, and Perplexity leaning on Reddit, LinkedIn and G2 for B2B questions (Search Engine Land).
For recommendation questions specifically it gets starker. Across 233 ChatGPT tool recommendations in 40 B2B SaaS categories, independent blogs and vendor content supplied 81.9% of cited sources (DerivateX).
This is the part most content plans miss. If the page an engine cites for "best X for Y" is a roundup on somebody else's domain, no amount of work on your own site changes that answer. The fix lives on the other domain: getting listed, getting the entry corrected, getting into the comparison.
Supporting this from a different angle, Ahrefs' analysis of 75,000 brands found that branded mentions across the web correlate with appearing in AI answers more strongly than domain authority or backlinks do (Ahrefs).
Not every visibility gap is a content gap. Some are PR gaps, and treating one as the other is how six months disappear.
What's the most common self-inflicted mistake?
Blocking the engines, usually by accident. It is boring, it is free to fix, and it is the one thing on this page that can take your visibility to zero on its own. [A]
Three versions of it, in order of how often we see them:
The nosnippet and max-snippet problem. Google's guide says pages need to be eligible to be shown with a snippet to appear in AI experiences. A nosnippet directive, or a restrictive max-snippet value, makes a page ineligible. Plenty of sites added these years ago to protect content from being read without a click, which is a perfectly reasonable thing to have wanted in 2019 and an expensive thing to still have in 2026.
Blocking AI crawlers in robots.txt. The bots do different jobs and blocking them has different consequences. Blocking a training crawler keeps you out of a future model's weights. Blocking a retrieval crawler like OAI-SearchBot, PerplexityBot or Claude-SearchBot keeps you out of answers being generated right now. A lot of sites blocked everything in one go during the 2024 backlash and never revisited it.
JavaScript-dependent content. Google renders JavaScript. The retrieval crawlers largely do not, and ChatGPT converts fetched pages to Markdown rather than rendering them. Content that only exists after a client-side fetch is content those engines cannot see.
All three are checkable in an afternoon. Our crawler access checker covers the first two if you want the quick version.
So what would we actually do first?
In order, weighted by evidence and by cost:
- Check you are not blocking anyone. Crawler access, snippet eligibility, JavaScript dependency. Free, grade A, and occasionally it is the whole answer.
- Find out where you actually stand. Not a one-off prompt in ChatGPT, which is a coin toss rather than a measurement, but a fixed question set run repeatedly so you have a denominator.
- Separate the content gaps from the third-party gaps. For every question you are losing, look at what is actually being cited. If it is a roundup you are not on, that is outreach, not a brief.
- Improve pages you already have before writing new ones. They are already crawled and may already be in the consideration set. Coverage gaps need new pages; quality gaps do not.
- Fix structure on the pages that matter. Standalone answers, question-shaped headings, tables for comparisons. Grade B evidence and a day's work.
- Then write, in clusters rather than one page per question. Engines split a question into sub-questions and assemble the answer from several sources, so covering a topic properly beats covering six topics thinly.
Notice that writing is sixth. That is not us being contrarian for the sake of it. It is where the evidence puts it, and it is why Answyn produces briefs against measured gaps rather than generating drafts against a calendar.
A note on the evidence
Everything above is graded, sourced and dated because this field has a citation problem of its own. A large share of published AEO research comes from companies selling AEO software, Answyn included. Sample sizes vary from four brands to 17 million URLs. Control groups are rare. Two respectable studies on the same question routinely disagree, and we have flagged it above where they do.
The grades are not a claim that grade A is true and grade D is false. They are a claim about how much of your budget each one deserves.
If you find a figure here that has moved, or a study we have read wrong, tell us at team@answyn.com and we will correct it and date the correction. This page is reviewed quarterly.
Questions people ask
- How do you get cited by ChatGPT?
- Be reachable, be relevant to the exact question, and be findable in the places ChatGPT already trusts. In practice that means checking you are not blocking OAI-SearchBot or suppressing snippets, having a page that answers the specific question in a passage that stands alone, and getting mentioned on the third-party sites that dominate recommendation answers. There is no ChatGPT-specific markup or file that helps.
- Does schema markup help AI citations?
- There is no good evidence that it does. A difference-in-differences study of 1,885 pages that added JSON-LD found no meaningful change on ChatGPT or Google AI Mode and a small decline on AI Overviews, and Google states directly that structured data is not required for generative AI search. Schema is still worth having for rich results and entity resolution, just not as a citation tactic.
- Does llms.txt work?
- No. Google says llms.txt files will neither harm nor help your visibility, and the most thorough synthesis of citation factors scored it lowest of 23 on the grounds that no credible experiment supports it. It costs twenty minutes to ship, so there is no harm in having one, but it should not be a paid line item.
- Do I need to rank in the top 10 to appear in AI Overviews?
- No. In Ahrefs' 2026 analysis of 863,000 keywords, 38% of cited pages ranked in the top 10, with roughly a third ranking 11 to 100 and another third ranking beyond 100 or not at all. Ranking helps and correlates strongly with citation, but it is not the gate.
- What is the single most common technical mistake in AEO?
- Accidentally blocking the engines. Either a nosnippet directive that makes a page ineligible for AI features, a robots.txt rule blocking retrieval crawlers that was added during the 2024 AI backlash and never revisited, or content that only loads via JavaScript. All three are free to fix and any of them can hold visibility at zero on its own.
- Is AEO just SEO with a new name?
- Largely, with a few genuine differences. Google's own guidance is that its AI features run on core Search ranking systems, so SEO fundamentals carry over. What is actually new is that answers get assembled from multiple sources per question, that a third of citations go to pages nowhere near the top 10, and that a large share of recommendation answers cite third-party sites rather than vendor sites at all.
- How long does it take to see results?
- Crawling happens within days, appearing in answers for long-tail questions takes weeks, and being consistently preferred over an incumbent takes months. The bigger issue is that single checks are unreliable, since the same question returns different sources on different days. Judge a page after four weeks of repeated measurement, not after one prompt.
- Why does my competitor get cited when their content is worse than mine?
- Usually because the citation is not about their content. Check what is actually being cited in the answer: if it is a third-party roundup, a Reddit thread or a review page that mentions them and not you, the gap is in earned mentions rather than in your writing. Branded mentions across the web correlate with AI visibility more strongly than backlinks or domain authority do.
Sources
- Google Search Central, AI optimisation guide
- Google Search Central, AI features and your website
- Zyppy, AI citation ranking factors (54 studies, 23 factors scored)
- Ahrefs via Search Engine Journal, AI Overview citations and ranking (863,000 keywords, 4m URLs, March 2026)
- Ahrefs, schema markup and AI citations (1,885 pages, difference-in-differences)
- Ahrefs, do AI assistants prefer fresh content (16.975m cited URLs)
- Ahrefs, AI Overview brand correlation (75,000 brands)
- Structural feature engineering for GEO, arXiv 2603.29979 (six engines)
- Seer Interactive, content recency and AI visibility (47,097 citations, 7,683 pages)
- Peec AI via Search Engine Land, most cited domains (30m sources)
- DerivateX, B2B SaaS AI citation study (233 recommendations, 40 categories)