Jack F, Co-founder
Now readingWhy it feels unanswerable
There is no universal number, and anyone giving you one is guessing. There is a sum you can run: half of it is measurement, half is stated assumptions, and this page shows how much the answer moves when the assumptions do. In the worked model below, moving from 6% to 20% unbranded AI visibility takes tens of pieces under any plausible assumption, and roughly 25 to 35 on the central case.
Why does "how much content" feel unanswerable?
It is one of the first questions anyone asks once a visibility score is on screen: how many pieces do we need to write before AI assistants like ChatGPT, Claude and Gemini start mentioning us? The answer most people get is "it depends", which is true and useless. It is what you say before anyone has put a denominator on the question.
Most people are running the numbers on a sample of one. Ask an assistant a category question, you are not in the answer, publish four pieces, ask again a fortnight later, still not there, conclude it is not working. That conclusion does not follow from the evidence, because the evidence was a single run on a single day, and a single run is a coin toss rather than a measurement.
The same question, put to the same engine on the same day, returns substantially different sources between runs: same-day source overlap sits at just 32 to 43% (University of St. Gallen, arXiv 2604.07585). So the fortnight-later check was never measuring the four pieces. It was measuring which way the noise happened to break that day.
The fix is not more content. It is a denominator large enough to average the noise out, which is what the rest of this page builds.
What is an AI visibility score actually made of?
Mentions divided by answers. Ask a fixed set of real customer questions, on a schedule, across the engines you track, and record two things for every answer that comes back: were you named, and was your site cited. Named in 288 of 4,800 answers is 6% visibility, which is just 288 divided by 4,800.
The moment you have a real numerator and denominator, "how much content does it take" stops being philosophical and becomes arithmetic.
One number from that same variance study is worth carrying into the sum: it puts the threshold for a stable per-question estimate at around seven runs. It is why the model below counts a full month of daily answers rather than a snapshot, and why the portfolio total, rather than any single question's rate, is what you should steer by.
How do you turn a visibility target into a number of pieces?
Here it is as a worked model. Say you are tracking 40 questions that do not mention your brand by name, the ones where somebody is describing a problem or asking who is good at something. Put those to four engines once a day for 30 days (four engines keeps the sum round, and the principle survives any roster) and each question produces 120 answers, so 4,800 answers in the window.
You are named in 288 of them. Six per cent.
You want to get to 20%. That is 960 mentions, so you need another 672.
Now, how many mentions does one question actually give you? If you cover a topic properly you still will not win every answer, because the engines disagree with each other, and with themselves between runs. Call it half. That is an assumption rather than a finding, it is the one with the most leverage in this sum, and we stress-test it in the next section. On it, winning a question outright is worth about 60 mentions.
672 divided by 60 is 11. You need to start winning roughly eleven questions you are currently losing.
Those eleven questions will not be scattered randomly. Tracked questions cluster into topics at around two to three per topic, so eleven questions is four or five topics. And covering a topic well enough to actually win it takes somewhere between five and eight pieces, for reasons we will come to.
Which, on the central assumptions, lands you here.
| Target visibility | Mentions needed | Questions to win | Topics | Pieces |
|---|---|---|---|---|
| 6% (today) | 288 | |||
| 12% | 576 | 5 | 2 | 10 to 16 |
| 20% | 960 | 11 | 4 to 5 | 25 to 35 |
| 30% | 1,440 | 19 | 7 to 8 | 40 to 55 |
So on the central assumptions, taking a brand from 6% to 20% costs about twenty-five to thirty-five pieces, on that prompt set, in that market. Push the assumptions to their plausible extremes and the range widens to roughly 15 to 55, which is still a useful answer, because it bounds the project to tens of pieces rather than a handful or hundreds. The sum also assumes your existing mentions hold while you add new wins, and that the questions you go after are ones where you are currently absent.
Run the same sum on your own numbers and you will get a different figure. That is fine. The point is that you get one, and you can plan against it.
How often do you win a question you've covered?
We assumed 50%, and you should not take that on trust, because it is the number with the most leverage in the sum. At a 30% win rate the 6% to 20% journey needs roughly 40 to 55 pieces. At 70% it needs 15 to 25. The answer is tens of pieces whichever way you lean, and it narrows to your number only when you measure it.
Here is the same journey at three win rates.
| If you win a covered question... | Questions to win | Topics | Pieces |
|---|---|---|---|
| 30% of the time | 19 | 7 to 8 | 40 to 55 |
| 50% of the time (central case) | 11 | 4 to 5 | 25 to 35 |
| 70% of the time | 8 | 3 | 15 to 25 |
The reason we still publish the sum with an assumed coefficient in it is that this is the one input you can both influence and measure. Once you are tracking, your real win rate is already sitting in the data: the share of answers that name you on the questions you have covered. Swap it in and the range above collapses to a plan.
Influencing it is where depth earns its keep. A topic covered with one page is a single bet on that page being retrieved. A topic covered with five or six pages that link to each other, define the same terms consistently and get cited together is a much better bet, and it lifts every page in the cluster, including the ones you had already published.
Two published findings explain why depth per topic beats one page per question. First, structure alone moves the number: a controlled experiment across six engines found structural optimisation by itself lifted citation rate by 17.3% (arXiv 2603.29979). Second, engines do not answer a question from one page. They split it into sub-queries and assemble the answer, which Google documents for AI Mode as query fan-out (Google), and only 38% of pages cited in AI Overviews rank in the organic top ten (Ahrefs, 863,000 keywords). You are rarely being outranked. You are being out-covered on the sub-questions. Five to eight pieces per topic is what covering the sub-questions actually looks like.
Depth does not change the maths. It improves one number inside it.
Should you write new content or update what you already have?
Updating is usually the cheaper win, and the evidence here is unusually clear. Seer Interactive analysed 7,683 pages carrying 47,097 citations across ChatGPT, Gemini and Perplexity. 75% of cited pages had been updated within the past year, and 88% within two (Seer Interactive).
Here is the finding that surprises people. Among pages where both dates were captured, 72% looked fresh by update date but only 42% by original publish date. The freshness the engines reward is being manufactured by maintenance, not by new publishing. And the pages cited consistently month after month were the older, maintained ones rather than the newest, so sustained visibility comes from established pages kept current.
So a rewritten page counts towards your number, and it usually costs less per point than a new one. If you have got a decent page that is not getting cited, improving its coverage and structure is often the cheaper win.
How long does content take to show up in AI answers?
Crawling happens within days of publishing, so a new page can be indexed almost immediately. Content becomes citable in retrieval-based answers for long-tail questions within a couple of weeks. Being consistently preferred over a competitor who already holds the answer takes months of maintained coverage.
That table above describes a quarter of work, not a fortnight. And because a defensible visibility number needs a full measurement window behind it, judging a piece after ten days does not tell you it failed. It tells you that you looked too early. Give a single piece four weeks and a topic cluster eight to twelve before you call it.
Can your own pages win every question?
No, and this is the part most content plans miss. For recommendation-style questions, the engines lean heavily on third-party editorial: across 233 ChatGPT tool recommendations in 40 B2B SaaS categories, independent blogs and vendor content supplied 81.9% of cited sources (Derivatex). When the cited page is a roundup on somebody else's domain, the fix lives on that domain too.
The wider evidence points the same way. Branded mentions across the web correlate with appearing in AI answers more strongly than domain authority or backlinks do (Ahrefs, 75,000 brands). So for some of the questions in your eleven, the "piece" your number needs is a placement or an updated entry on a page you do not own, alongside the cluster you are building at home. Same arithmetic, different address. Not every visibility gap is a content gap, and we have written separately about the other kinds.
Should you focus on one topic, or chase the gaps?
Both, in the right order: use the gap data to choose which topics to build, then go deep on two or three of them. Chasing gaps one question at a time leaves you with twelve lone pages across six topics, each relying on being retrieved entirely on its own merits. That is the low win rate scenario from the table, produced by exactly the same amount of work.
Rank the topics by how many tracked questions sit in each, how much room there is in the answers, and how much you actually want the customers asking them. Then saturate. Then move to the next one.
Same data, same instinct, better unit.
How do you run this sum on your own numbers?
Count your unbranded questions, count the answers collected in your window, count your mentions, pick a target, and work backwards. Tracking hands you the measured half of the sum on day one: your visibility and your mentions gap. Your covered-question win rate follows within weeks. Pieces per topic is the one input that stays a budgeting estimate.
If you are not tracking it, that is the actual first job, because everything above depends on having a denominator. Without one you are back to "it depends", and you will be there for a while.
This sum is also what Answyn's briefs are built around. We show mentions over answers for every tracked question, cluster the losses into topics, and turn each one into a brief rather than a generated draft. Pro includes 30 a month and Business includes 100, which at five to eight pieces per topic is roughly one topic cluster a month. And because per-question numbers wobble, we would rather you judged movement over a full window than off a single run.
Questions people ask
- How many pieces of content does it take to show up in AI answers?
- There is no universal number, but there is a sum you can run for your brand: count the extra mentions you need to hit your target, then divide by the share of each question's answers you expect to win. In a worked model of a typical set (40 unbranded questions, four engines, daily answers for a month), going from 6% to 20% visibility took 25 to 35 pieces on central assumptions, and 15 to 55 across plausible win rates.
- Is it better to update existing content or publish new content?
- Updating is usually the cheaper win. Seer Interactive's analysis of 7,683 cited pages found 75% had been updated within the past year, and pages looked far fresher by update date (72%) than by original publish date (42%). If you have got a decent page that is not being cited, improve it before you write a new one.
- How long does it take for new content to appear in AI answers?
- Crawling happens within days. Appearing in retrieval-based answers for long-tail questions typically takes weeks. Being consistently preferred over a competitor takes months. Do not judge a single piece before four weeks, or a topic cluster before eight to twelve.
- Should I focus on one topic or cover all my visibility gaps at once?
- Use the gap data to choose which topics to build, then go deep on two or three. Writing one page per gap leaves you with scattered pages that each have to win retrieval on their own, which is the lowest-probability version of exactly the same amount of work.
- Can AI visibility actually be measured, or is it guesswork?
- It is measurable as long as you are recording it question by question, with enough answers behind the number. Ask a fixed set of customer questions across the engines on a schedule, note whether you were named and whether you were cited, and your visibility is mentions divided by answers. Single runs are noisy, so the published threshold for a stable per-question estimate is around seven runs per engine.
- What counts as a good AI visibility score?
- That depends far more on your category than on any universal benchmark, because the number of brands an answer names varies enormously by sector. The figure worth watching is your unbranded score, measured on questions that do not contain your brand name, because that is where new customers actually find you.
- Do I need different content for each AI engine?
- You do not need engine-specific versions of a page: the things that make content citable, meaning coverage, structure and freshness, are common to all of them. But the engines differ in what they retrieve, and cross-engine overlap in cited sources is low, so expect uneven movement rather than a neat straight line, and measure each engine separately.
Sources
- Seer Interactive, content recency and AI visibility, 47,097 citations across 7,683 pages
- University of St. Gallen, Don't Measure Once, arXiv 2604.07585
- Structural feature engineering for GEO, arXiv 2603.29979
- Ahrefs via Search Engine Journal, 863,000 keywords
- Ahrefs, AI Overview brand correlation, 75,000 brands
- Derivatex, B2B SaaS AI citation study, 233 recommendations
- Google, AI features and your website
The worked example is a planning model, not a measurement from a tracked brand. The measured inputs are yours to supply. The assumptions are stated and stress-tested above, and the sum works with any of them swapped for observed rates. Answyn is a product of Luto Ventures Ltd, company 16563350, registered in England and Wales. This page is reviewed quarterly. If you find a figure that has moved, tell us at team@answyn.com and we will correct it and date the correction.