Does ChatGPT Give Everyone a Different Answer? Measuring AI Visibility When Every Reply Is Personalised
The sharpest objection to AI visibility tracking: if the model personalises every answer, isn't measuring it through a tool measuring a fiction? Here is the honest answer. Personalisation re-ranks the shortlist, it does not write it, and the signal you can actually influence is the shared, un-personalised one a good tracker measures.
Here is an objection worth taking seriously, because the sharpest people in search raise it. When you ask ChatGPT a question, it does not answer a generic you. It answers the specific you: your past chats, your saved memory, your custom instructions, roughly where you are. So the argument goes: any tool that measures AI visibility through an interface rather than a real logged-in session is measuring a fiction, because it never sees the context the model has about the person asking. The objection is half right, and the half it gets right is the half that does not matter for the decision you are trying to make. This post is about which half is which, and why we measure the way we do.
Start by conceding the part that is true
The fastest way to lose a smart audience is to wave away a real objection, so let us state the strong version plainly and agree with it. When a real person asks ChatGPT for a good accountant, an electrician who can come Tuesday, or somewhere to eat near the office, the model is not reasoning from a blank slate. It may know the person has asked about small-business tax before. It may hold a saved instruction to prefer independent firms. It knows, roughly, where they are. Two people can type the same words and get different names back. That is real, it is by design, and it is getting more true over time, not less.
So the first thing to say to anyone who raises this is: yes. No tool can show you the answer a specific, named individual sees inside their own account, because that answer is private, personalised, and different for everyone. Here is the part the objection skips over: neither can a human. If you log into your own ChatGPT and check, you are looking at exactly one person's personalised answer, your own, shaped by your own history, impossible to reproduce tomorrow, and useless as a benchmark for anybody else. The choice was never between a clean interface measurement and the real personalised truth. There is no measurable personalised truth. The choice is between a stable, repeatable baseline and no measurement at all.
Concede the boundary cleanly, then show it does not touch the decision. You cannot measure what one stranger sees. You can measure the thing that decides whether you are in the running for all of them.
Personalisation re-ranks the shortlist. It does not write it.
This is the load-bearing idea, so it is worth being precise. When an answer engine handles a local question, it first assembles a set of candidate businesses it considers plausible, drawn from what it learned in training and, increasingly, from a live search it runs on your behalf. Then, if it has anything personal to go on, it leans the ordering of that set towards what suits this particular person. Personal context is a re-ranking signal applied late, on top of a shortlist that already exists.
The consequence is simple and it is the whole ballgame. If your business is not in the candidate set, no amount of personalisation can promote you into it. You cannot be re-ranked into an answer you were absent from. So the question a visibility measure needs to answer is not “what order does one person see”, it is “are you in the set the engine is willing to draw from at all”. That property is far more stable than the personalised ordering laid on top of it, because it is a function of the model's knowledge and the sources it retrieves, not of who happens to be asking.
The personalisation that matters most helps people who already know you
Here is the part that quietly defuses the objection on commercial grounds. The strongest personalisation signals, saved memory and past conversation history, bias the model towards things the person has already engaged with. If someone has talked to ChatGPT about your brand, or a category you sit in, the model is more likely to echo that back. That is a real effect. It is also, almost entirely, a retention effect. It rewards familiarity you have already earned.
The visibility that actually grows a business is the opposite case. It is the new prospect who has never heard of you, has no relevant history, and asks a cold question into an assistant that knows nothing about their preferences in your category. That moment, the one that decides whether you win a customer you did not already have, is by definition the least personalised query there is. An anonymous, standardised measurement models that acquisition moment better than a logged-in test ever could, because a logged-in test is contaminated by whatever the tester already knows and has already searched.
What personalisation mostly changes
- •The ordering for someone who already knows your brand or category.
- •Answers for a person whose saved memory or custom instructions nudge a preference.
- •A single individual's screen, which nobody else can reproduce or benchmark against.
- •Retention: reminding a warm audience of a name they already carry.
What actually decides new-customer visibility
- •Whether you make the candidate set a cold prospect's question draws from.
- •The sources the engine retrieves and cites, which are the same for everyone.
- •Your share of voice against competitors under identical, neutral conditions.
- •Acquisition: getting named to someone who did not arrive already knowing you.
You can only optimise the un-personalised layer anyway
Suppose we could see inside a stranger's personalised session. What would you do with it? You cannot edit their saved memory. You cannot rewrite their custom instructions or change what they searched last month. The private context that personalises their answer is, from your side of the glass, completely inert. There is no lever there to pull.
The levers that do exist all live in the un-personalised layer, which is exactly the layer a standardised measurement exposes. Whether the engine can tell precisely who you are (entity clarity). Whether it can lift a clean answer straight off your pages (extractability). Whether the wider web agrees about you strongly enough that recommending you feels safe (consensus). Move those and you shift the shortlist for everyone, warm and cold alike. So the surface a good visibility tool measures and the surface you can actually change are the same surface. Measuring the personalised re-ranking, even if it were possible, would be measuring the one thing you have no way to act on.
You cannot optimise a stranger's memory. You can optimise whether the machine is confident enough to say your name at all. A visibility measure earns its keep by tracking the second thing, not chasing the first.
The context that is commercially real is not lost, it is controllable
The strongest version of the objection usually retreats to one specific piece of context: location. “Near me” is the whole game in local search, and if a measurement ignores geography it is measuring the wrong thing. Agreed. The answer is that location is not some private, unknowable attribute of the person. It is a parameter you can set deliberately.
So rather than inheriting one tester's accidental location, a measurement can fix the location on purpose, to the market you care about, and hold it steady across every check. That is better than a real logged-in session, not worse, because it is intentional and repeatable. The same is true of the live web results the engine grounds its answer in: those are shared, public and observable, and the citations that come back are the single most objective piece of the whole picture. When your measurement reports which sources the engine pulled to build its answer, you are looking at evidence that is completely independent of who was asking.
The honest way to measure a noisy surface
The deepest reason the objection feels powerful is that it is really an objection to a specific, bad way of measuring: taking one answer, at one moment, and treating it as a verdict. On a surface this volatile that is genuinely unsound, and not only because of personalisation. Ask the same question twice, minutes apart, logged out, and the wording, the names and the order can shift. One answer is a coin toss, not a measurement.
The fix is not a more “real” single answer. It is to stop relying on single answers. A sound measurement treats each question as a sample and reports a distribution:
- 1
Run each question more than once
A single reply is one draw from a noisy process. Running the same question several times and reporting how often you appear turns a coin toss into a frequency you can trust. “Named in most checks” is a real finding. “Named that one time I looked” is not. - 2
Fix the context on purpose
Set the location to the market you care about and hold it constant, so a change you see is a change in the world, not a change in where the tester happened to be. Do not inherit one person's accidental context, choose the context deliberately. - 3
Measure relative, not absolute
Report share of voice against the competitors who keep recurring, not a made-up “rank”. Personalisation and noise hit an absolute position hard. They mostly wash out of a relative comparison run under identical conditions. - 4
Read the trend, not the snapshot
Noise is roughly constant, so it cancels over time. The line that matters is whether your visibility rate and share of voice climb as you strengthen your signals. That trend is what personalisation cannot fake and cannot hide.
Put those together and the objection loses its force. Personalisation perturbs any single observation. It does not perturb the aggregate of many observations, run under a fixed context, compared against the same competitors, tracked over weeks. That is not a workaround for a weakness. It is simply how you measure anything on a probabilistic surface, and it is the same discipline that has always underpinned rank tracking, which has been personalised, localised and device-dependent for well over a decade.
The objection, followed all the way down, proves too much
There is one more tell worth naming. If “you cannot see the personalised answer, so the measurement is meaningless” were actually true, it would not just sink one tool. It would sink the entire category, and Google rank tracking with it, since Google results have been personalised and localised for years. Yet brands demonstrably gain and lose visibility on both surfaces, and they do it by moving the observable, optimisable factors. An argument that proves no measurement can ever exist is not a rigorous standard. It is a counsel of despair dressed up as sophistication, and the practical effect of accepting it is to fly blind while your competitors do not.
How we actually measure it, and what we deliberately do not claim
For the record, this is how we approach it, because naming the objection and then living up to the answer is the only version that earns trust. We never claim an AI “rank”, because none exists. We turn the keywords you already track into the real, contextual questions a customer would ask, set the location to your market, and report three things: your visibility rate (how often you show up), your share of voice against the competitors who keep appearing, and the sources the engine cited to build each answer. We save every reply word for word, so nothing is a number you have to take on faith. And we treat the trend over time as the headline, not any single check.
What we will not tell you is what one specific, named person sees in their own logged-in session tonight. Nobody can tell you that honestly, and a vendor who claims to is the one you should distrust. What we can tell you, reliably and repeatedly, is whether you are in the running when the questions your customers ask get asked, and whether the work you are doing is moving that in your favour. That is the number that changes when your marketing works, and it is the one worth managing.
Keep reading