Jack F, Co-founder
Part of the Answyn AEO Handbook.
Now readingWhat it is
Customer journey analysis in AI search means measuring how often AI assistants mention your brand at each stage of the customer journey, rather than as one overall score. You tag each question you track with a journey stage, then read your visibility stage by stage, which shows you where in the decision the conversation stops carrying you.
Buyers have not changed how they decide. They still work out they have got a problem, look at what is available, compare a shortlist and check their favourite before committing.
What has changed is that an AI assistant now sits in the middle of all four of those moments, and it names brands at wildly different rates depending on which moment you are in. In one published study, brands were named in 0.10% of answers to early problem-shaped questions and 1.29% of answers to "what are the best X" questions. Later on, when someone asks about a brand directly, that jumps to 96%.
So if you are measuring one visibility number across everything, you are averaging together stages that behave nothing like each other. This piece walks through what happens at each one, and what to do about it.
What is customer journey analysis in AI search?
Customer journey analysis in AI search means measuring how often AI assistants mention your brand at each stage of the customer journey, rather than as one overall score. You tag each question you track with a journey stage, then read your visibility stage by stage, which shows you where in the decision the conversation stops carrying you.
If you are new to this, the mechanics are simpler than they sound. You write out the questions your buyers actually ask, in their words. That is your prompt set. You run those questions past the assistants your buyers use, repeatedly, and record whether you got mentioned. Tagging each question with a journey stage is the only extra step, and it is the one that turns a number into a diagnosis.
Without it, a visibility score of 40% could mean you are evenly present all the way through, or it could mean you are at 5% while people are working out what they need and 75% once they have typed your name. Those are two completely different businesses with the same number on the dashboard.
What are the stages, and what does a buyer actually ask?
Six stages, in five bands. The band tells you who is asking. The stage tells you what they asked.
| Band | Stage | Who's asking | A question at this stage sounds like |
|---|---|---|---|
| Awareness | Problem | Someone with a symptom | our board keeps asking about our visibility in AI search and I have no idea how to measure it |
| Consideration | Options | A warm lead | what tools track this sort of thing, and which are worth a look |
| Consideration | Comparison | A warm lead | how does A compare to B if we're a small team |
| Convert | Validation | Someone deciding | is A any good, what do people say about them |
| Loyalty | Retention | An existing customer | is A still worth what we pay, should we switch |
| Advocacy | Advocacy | A happy customer | would you recommend A to someone in our position |
Most models fold options and comparison together into one consideration stage. It is worth keeping them apart, because they behave very differently, and we will come back to why further down.
That first example is a real question from a tracked prompt set. Notice what it looks like: long, conversational, describing a situation rather than naming a category. You do not arrive at a question like that through keyword research, which is the first practical thing to take from all this.
Problem: you're probably not there
This is where the drop-off is sharpest. Someone describes a symptom and gets a synthesised answer about the symptom, which pulls in half a dozen sources and names almost nobody.
Two useful numbers. Brand mentions on problem-awareness questions ran at 0.10%, against 1.29% on category-research questions (Victorious Quarterly Search Report Q2 2026, 150 brands, 5 verticals, 8 platforms, 5,830 responses). And across AI answers generally, 61.7% of citations do not name a brand at all (Semrush and Kevin Indig, Ghost Citations, 3,981 domain appearances across 115 prompts, 14 countries and 4 engines, 9 June 2026).
That second figure is worth sitting with, because it catches people out. Getting cited and getting mentioned are two different outcomes. Your page can be one of the sources an assistant used to build its answer, while your name appears nowhere in the answer the buyer reads. At the problem stage that is the normal result rather than the exception.
The upside is that nobody else is winning here either. It is an open field rather than a crowded one.
Options: this is where the shortlist forms, and it's now the biggest influence there is
The old version of this was a listicle of ten that a buyer picked from. Now the assistant is the listicle, and it typically names somewhere between two and six brands.
That short list matters enormously. AI chatbots are now the single biggest influence on B2B shortlists, at 54%, ahead of software review sites at 43% and analyst firms at 36%. 69% of buyers ended up choosing a different vendor from the one they had planned on, and 33% bought from a vendor they had never previously heard of (G2, The Answer Economy, 1,076 B2B decision-makers, March 2026).
Read that last one again if you are new to this. A third of buyers are ending up with a company they had never encountered before an assistant introduced them.
Comparison: the most stable ground you've got
Ask an assistant to compare two named brands and something different happens. Comparative questions carry the highest brand mention rate of any question type, at 43.3%, against 18% for informational ones (Semrush, Ghost Citations).
They are also the most consistent. BrightEdge found 80% agreement between engines on "compare" questions, against 23% on "best" questions (27 August 2025, so a year old now, but it points the same way as everything else we have seen).
High agreement is good news, because it means a fix here tends to hold across engines rather than needing doing four times. If you only have budget for one thing, comparison pages are usually it.
Validation: your best numbers, and your least useful ones
By this point the buyer has typed your name, so the assistant only has to find your website. Almost everyone looks good here.
The clearest illustration: 96% of brands were described accurately when an assistant was asked about them directly, while 89% never appeared in category-research answers at all (Victorious Q2 2026). Recognition is close to universal. Recommendation is close to zero.
Which means a visibility score built mostly from branded questions will flatter you and tell you very little. It is an easy trap to fall into by accident, because branded questions are the obvious ones to think of first.
You can still lose here, mind. 64% of B2B buyers encounter inaccurate AI recommendations often or very often, and when an assistant leaves out a brand a buyer already trusts, 21% of them go with the assistant's suggestion anyway (G2).
Retention and advocacy: nobody has measured these
We went looking for any published measurement of what AI assistants say to a brand's existing customers, from any vendor, agency or research group, and found nothing.
The questions are clearly being asked. "Is this still worth what we pay." "Should we cancel." "Is there something better than what we have got." Those go into a system with no loyalty to you and perfect recall of your competitors. G2 found buyers are still using ChatGPT at the retention stage, at 57% of them, so we know they are in there. We just do not know what they are being told.
Advocacy has an odd wrinkle too. 45% of B2B buyers say review-site citations are the single most confidence-inspiring thing they can see in an AI answer. But in the only controlled study of B2B software recommendations we could find, G2 and Capterra received zero citations across all 233 recommendations, with 81.9% coming from independent blogs and vendor content instead (DerivateX, 40 categories, 219 tools, 31 May 2026). Buyers are asking for a signal the assistants do not appear to be showing them. That is one study in one category, so hold it lightly, but it is worth knowing before you spend heavily on review platforms.
The thing that surprised us most
Here is the one place our own data adds something you cannot get elsewhere.
We ran 50 stage-tagged questions past four assistants for 30 days and looked at how far apart their answers were. Not how well one brand did, but how much the engines disagreed with each other about the same question on the same day.
| Stage | How far apart the four engines were |
|---|---|
| Problem | 2 points |
| Options | 50 points |
| Comparison | 8 points |
| Validation | 8 points |
Three stages out of four, the assistants broadly agree. At the options stage they do not agree at all. On the same questions in the same window, ChatGPT named the brand in 53% of answers and Gemini named it in 3%. Google AI Mode came in at 18%, AI Overviews at 23%.
That is a problem, because options is the stage that forms the shortlist, which we have just established is the biggest single influence on the decision. The number you most need is the number a blended score is least able to give you honestly. A single figure sitting between 53% and 3% describes neither.
It is also the reason we keep options and comparison as separate stages. They produce similar-looking visibility figures, and they behave nothing alike underneath.
What to do with that: read your visibility per engine, at least at the options stage. If you only ever look at one blended number, you cannot tell whether you have got a broad problem or one engine that has never heard of you.
Reading the shape of your curve
Once you have got visibility by stage, the useful thing is the shape rather than any single number. A low stage is not automatically a problem, because near-zero at the problem stage is what almost everyone gets. A stage where visibility falls is always worth a look, because it means something specific went wrong at a specific moment.
Four shapes come up again and again:
| Shape | What it looks like | What it usually means |
|---|---|---|
| Conversion gap | Visibility drops between options and comparison | You are named on the shortlist and missing from the head-to-head. Normally comes down to missing or out-of-date comparison pages. |
| Discovery gap | Near-zero at problem, healthy everywhere else | People who already know you find you. People with the problem never meet you. |
| Invisible positioning | Present throughout, but the descriptions are vague or wrong | An accuracy and entity problem rather than a content one. |
| Category default | Strong everywhere, suspiciously flat | Usually a sign your question set is mostly branded, so rebuild that before drawing conclusions. |
The conversion gap is the most common of the four, and it is the one worth checking first.
Two things to hold on to when you look at yours. Check it per engine, because a fall on one assistant gets cancelled out by a rise on another and the blended view will show you a smooth line that exists nowhere. And only read the pre-purchase stages, because retention and advocacy are a different person asking a different thing.
How to start
If you are setting this up for the first time, five things in order:
- Write 30 to 50 questions in your buyers' actual words. Not keywords. Full sentences, the way someone would type them into ChatGPT at half past four on a Tuesday. Sales calls and support tickets are the best source.
- Tag each one with a stage before you collect anything. You cannot add a stage retrospectively to a window you have already run, and untagged questions cannot sit on the map.
- Balance them across stages. Six or eight per stage beats fifty at the bottom. A set that is mostly branded questions will make you look great and teach you nothing.
- Include retention and advocacy questions anyway. There is no benchmark to compare against yet. There will not be one until people start collecting.
- Read the results per engine, and per stage. Two dimensions, and both matter. One blended number hides both.
That is the whole method. The hard part is step one, and it is the part worth doing properly. If you would rather not assemble it yourself, Answyn's Customer Journey does the stage mapping and the per-engine split as a standard view.
A note on our data
The cross-engine figures in this piece come from Tilio, a UK AEO agency and a business connected to Answyn. 50 tracked questions, tagged by stage, run against ChatGPT, Google Gemini, Google AI Mode and Google AI Overviews in the UK over a rolling 30-day window ending 27 August 2026, giving 1,300 answers. Claude, Perplexity, Microsoft Copilot and Meta AI were not in this window.
It is one brand in one market, so treat the shape as the interesting part rather than the levels. Everything else in this piece is cited to its original source below.
Questions people ask
- What is the customer journey in AI search?
- The same sequence of decisions buyers have always made, asked as conversational questions to an AI assistant instead of typed as keywords into a search engine. Six stages: problem, options, comparison, validation, retention and advocacy. What has changed is that a synthesised answer now sits between the buyer and your website at every one of them.
- Is the marketing funnel dead in AI search?
- The stages are all still there and buyers still move through them in roughly that order. What has changed is where brands get named, which is heavily weighted towards the later stages. The assistants are largely adjudicating demand that already exists rather than creating it.
- Do different AI engines show the same customer journey?
- No, and the disagreement is concentrated in one place. In our 30-day window the four engines landed within about eight points of each other at the problem, comparison and validation stages, and fifty points apart at the options stage. That is the case for reading visibility per engine.
- Which journey stage matters most?
- Options, on the evidence. It is where the shortlist forms, and AI chatbots are now the biggest single influence on B2B shortlists at 54% (G2, 1,076 buyers, March 2026). It is also the hardest stage to measure reliably, which makes it easy to be wrong about.
- Why does my brand show up in ChatGPT for branded questions but not category ones?
- Because those questions ask different things of the assistant. A branded question only needs it to find your website, which it can nearly always do. A category question needs it to pick you out of a field. One study found 96% of brands described accurately when asked about directly, while 89% never appeared in category answers (Victorious Q2 2026).
- Can you measure AI visibility after someone has bought?
- In principle, by tracking retention and advocacy questions. In practice almost nobody does, and we could find no published measurement of what assistants say to existing customers. Worth starting now, because you cannot go back and collect a window you did not run.
- How many questions do I need to track?
- Fewer than most people expect, spread better than most people manage. Thirty to fifty across six stages is a workable start. Balance across the stages matters more than the total.
- Can I do this without a tool?
- For a one-off snapshot, yes. Write your questions, tag them, ask each assistant by hand and record what comes back in a spreadsheet. It falls apart at scale, because the same question returns different brands on different days, so a single run tells you very little. You need each question asked repeatedly, across several assistants, over weeks.
- Which tools can track visibility by journey stage?
- Not many, at the time of writing. HubSpot's AEO product has a journey-stage filter. Peec AI lets you build one using custom tags, though there is no built-in stage taxonomy. Answyn has Customer Journey as a standard page with all six stages and the curve shapes above. Most other AI visibility platforms group by topic or intent rather than by journey stage, so you would be tagging manually.
- How much does journey tracking cost?
- Answyn starts at £149 a month for 50 tracked questions across 2 engines daily, and £349 for 200 questions across 4. Billed in pounds, which is unusual in this category, as most platforms price in dollars or euros. Current details are on the pricing page.
Sources
- G2, The Answer Economy (1,076 B2B decision-makers, March 2026)
- Victorious Quarterly Search Report Q2 2026 (150 brands, 8 platforms, 5,830 responses)
- Semrush and Kevin Indig, Ghost Citations (3,981 domain appearances, 9 June 2026)
- DerivateX, B2B SaaS AI citation study (233 recommendations, 31 May 2026)
- BrightEdge, cross-engine disagreement (27 August 2025)
The cross-engine figures in this piece come from Tilio, a UK AEO agency and a business connected to Answyn.