Which AI Prompts Should a Local Business Track? Build Coverage, Not a Wish List
Most AI visibility tracking fails at the first step, with a list of prompts written from memory in a meeting. Here is what replaces it. Take the keywords you already track and the wording customers already give you in calls, enquiry forms and reviews, expand each into a short ladder of buying-intent questions, and let the shape of that ladder change by sector. Includes why coverage beats guesswork, how to read the results as a consensus problem rather than a wording problem, and worked examples for emergency trades, dentists, restaurants, builders and professional firms.
Ask a local business owner which prompts they want to track in ChatGPT and you will have a list inside ninety seconds. Best plumber in Leeds. Top rated dentist Manchester. Cheapest accountant near me. It is a fast, confident list, and it is the wrong shape, because it describes the answers the owner would like to win rather than the range of ways a customer arrives at the question. Nobody can know the exact sentence a stranger typed last Tuesday. What you can do is cover the space of reasons people ask, and then find out whether the wider web backs you up across it. This post is about how to build that coverage: where the raw material comes from, what shape it takes, why it is different for an emergency plumber than for an orthodontist, and what the results are actually telling you.
Start with what people actually type
This is no longer a niche behaviour you can reason about from your own habits. About half of US adults now report using AI chatbots, up from a third two years earlier.1 More to the point, the largest study of what they use them for, run over a representative sample of real ChatGPT conversations, found that Practical Guidance and Seeking Information sit alongside Writing as the three dominant categories, together accounting for close to eighty per cent of conversations.2 Practical guidance is exactly the shape of who should I call about this.
~50%
of US adults use AI chatbots
Up from a third in 2024, on a survey of 5,119 adults.
~80%
of conversations are guidance, information or writing
Classified across a representative sample of real conversations.
2 to 3
words in a typical search query
The habit an assistant question is not constrained by.
Now put that next to how people search. Decades of usability research put the typical search query at two or three words, reaching for the familiar term rather than the clever one.3 A question to an assistant has no such ceiling, so it carries the things a search box forced people to drop: the situation, the deadline, the budget, the constraint. When researchers wanted to study real assistant use they had to go and collect a million consented conversations, precisely because prompts written by researchers do not resemble what people actually send.4 That is the same trap as the ninety-second list. A list written from your own head is a list of what you think the question looks like.
Measures almost nothing
- •best plumber Leeds
- •plumbing services near me
- •emergency plumber
- •who is the best plumber
Measures something
- •my boiler is leaking and I need someone in Leeds tonight
- •find me an emergency plumber in Leeds who can come out now
- •who do I call in Leeds if a pipe bursts at the weekend
- •what does an emergency plumber in Leeds charge for a night callout
The left column is not a strawman, it is what most tracking lists look like on day one. Each entry is a keyword wearing the costume of a question. They have no situation in them, no constraint, and two of them have no place at all, so an assistant either hedges, answers about the wrong city, or produces a general essay about choosing a plumber. Meanwhile a superlative invites exactly the response you cannot use: it depends what you need, here are some things to look for. You run a list like that, find yourself missing from most of it, and conclude you have a visibility problem. You may well have one. What you definitely have is a coverage problem, and until that is fixed you cannot see past it.
Where the raw material comes from
You already own everything you need to write the right-hand column. It arrives in two streams, and the second one is the one almost nobody uses.
The keywords you already track
If you run local rank tracking, that list has already been filtered by reality: somebody picked those terms because they bring work in, they have survived contact with your Search Console data, and the services behind them are ones you want more of. Keep the ones with money behind them and cut the purely informational. What remains are the topics, and every topic becomes a small ladder of questions rather than a single line.
Your own first-party record of how customers talk
This is the good stuff, and it is sitting in systems you already pay for. Your customers have been describing their problem to you, in their own words, for years. Nobody has to invent anything: it needs collecting and tidying, not imagining.
Calls and reception notes
The first sentence of an enquiry call is almost always a perfectly formed query. Call recordings, transcripts, or just the notes your receptionist types while the customer is still talking.
Enquiry forms and quote requests
The free-text box on your contact form is unfiltered customer language. So is the tell us about your job field, and the emails people send before they book.
Live chat and message threads
Chat logs, WhatsApp enquiries and the questions that arrive through your Google Business Profile. These are already typed, already conversational, already in the format you need.
Reviews and complaint themes
Reviews explain what mattered after the fact, which tells you the deciding factor: parking, punctuality, whether someone explained the price before starting.
The conversion is mechanical. Take what the customer said, drop the pleasantries, keep the constraint, and write it as one sentence a stranger could send cold. The constraint is the valuable part and the part an invented list never contains, because you do not think to invent before Friday, or with parking, or who takes NHS referrals, or someone who has done a listed building before.
What the customer said Verbatim, from a call or a form | The query it becomes Same meaning, sendable cold | |
|---|---|---|
| Dental | Hi, my daughter is 14 and we have been told she needs braces, we are in Chorlton and can only really do after school | Which Manchester orthodontists treat teenagers and offer after-school appointments |
| Trades | We are getting quotes for a two storey extension, the last builder said planning would be a problem | Who handles planning and building control for two storey extensions in Guildford |
| Professional | Our accountant has retired, we sell on Shopify into Europe and nobody seems to understand the VAT | Which accountants in Reading deal with VAT for ecommerce businesses selling into the EU |
Why coverage, and not guesswork
Here is the part that makes the method make sense, and it is the reason we draft questions for you rather than handing you an empty box.
When someone asks an answer engine a real question, the engine does not run that one sentence and stop. Google describes AI Mode as using a query fan-out technique, issuing multiple related searches concurrently across subtopics and multiple data sources, then bringing the results together into one response.5 ChatGPT does its own version, turning your request into searches of its own and assembling a reply from what comes back.6 We have a full guide to query fan-out because the implication is large: the sentence a customer types is the start of a retrieval process, not the whole of it. Tracking one phrasing tests one narrow slice of that process.
A tracked query is not a keyword you rank for. It is a probe into a space, and coverage is how many parts of that space you have probed.
The second reason is what you are ultimately measuring. Research on how models learn found a strong relationship between a model's accuracy on a fact and how many documents supporting that fact appeared in its training data.7 Things the web says often are known well; things the web barely mentions are known badly, and scaling the model up helps far less than you would hope. Retrieving pages at answer time softens this but does not remove it,8 because retrieval still has to find something to bring back. For a local business this translates cleanly. Your visibility in an answer is downstream of what the wider web says about you, how consistently it says it, and in how many places. That is the same prominence idea Google has documented for local results for years.9
Which is why the point of a coverage plan is not to guess wording. It is to probe that consensus from enough angles to find where it is thin. If you are named confidently for one service and absent on the neighbouring one, you have not discovered a prompt-writing failure. You have found the edge of your own footprint, and the fix lives in entity clarity, citation consistency and the rest of your local signals rather than in the question you asked.
One keyword, a ladder of queries
Every keyword can be asked from several distances. Someone standing in a flooded kitchen asks differently from someone with three quotes on the table, who asks differently again from someone idly wondering what a job like this costs. Those are rungs on a ladder, and how many rungs you take is a budget decision, so it is worth knowing what each buys you.
- 1
Ready to buy
2 questions per keywordBoth rungs sit at the point of decision. This is the default, and it should be, because these are the questions where answer engines name specific businesses at all. If you only ever run this, you still have a real measurement.
- 2
Comparing
3 questions per keywordOne decision anchor, then the ways people choose between the names they were given. Worth the extra rung when you already appear sometimes and want to know what you are losing on: price, credentials, reassurance.
- 3
Full funnel
4 questions per keywordDecision, comparison, and exactly one earlier question. One, never more. Earlier questions are where engines are least likely to name anybody, so that slot only earns its keep once the buying end is covered.
The ladder changes shape by sector, and that is the whole trick
Here is where generic advice falls over. The rungs are not the same for every business, because the reasons people ask are not the same. What sits behind a question is what we call the angle: urgency, price, credentials, occasion, whether you can start this year. A dentist and an emergency plumber both have a decision rung and a comparison rung, but the contents have almost nothing in common, and a plan that treats them identically ends up measuring the wrong half of one of them.
Three worked examples, then two more in brief. Read them for the pattern rather than the words: what changes between them is the order of the angles, and the order of the angles is what your sector decides.
Emergency trade
Tracked keyword
emergency plumber leeds
“my boiler is leaking and I need someone in Leeds tonight”
Urgency
“find me an emergency plumber in Leeds who can come out now”
Imperative
“who do I call in Leeds if a pipe bursts at the weekend”
Situation
“what does an emergency plumber in Leeds charge for a night callout”
Callout price
“which emergency plumbers in Leeds guarantee their work”
Guarantee
Clinical and considered
Tracked keyword
invisalign manchester
“who is good for Invisalign in Manchester”
Recommendation
“which dentists in Manchester could start Invisalign within a month”
Booking
“Invisalign in Manchester with evening or Saturday appointments”
Constraint
“which Manchester Invisalign providers have done the most cases”
Credentials
“how much does Invisalign cost in Manchester and who offers payment plans”
Price
Hospitality and retail
Tracked keyword
italian restaurant bristol
“where is good for Italian food in Bristol”
Recommendation
“somewhere Italian in Bristol for an anniversary dinner”
Occasion
“book me a table at an Italian restaurant in Bristol on Saturday night”
Imperative
“which Italian restaurants in Bristol handle gluten free properly”
Suitability
“which independent Italian places in Bristol are better than the chains for a quiet meal”
Comparison
Two more, compressed. The skeleton is the same, three rungs at the point of decision, but watch what happens at the comparison stage: for a builder the real second question is quotes and risk, and for a professional firm it is whether you have done this for someone like them.
Trades project Extensions, roofing, landscaping | B2B professional Accountants, IT support, HR | |
|---|---|---|
| Keyword | builder guildford | accountant reading |
| Rung 1, decision | Who is good in Guildford for a two storey extension | Which accountants in Reading work with small ecommerce businesses |
| Rung 2, decision | Find me a builder in Guildford who could start an extension this year | Find me an accountant in Reading who can take over our year end |
| Rung 3, decision | Who handles planning and building control for extensions in Guildford | Which Reading accountants deal with VAT on sales into the EU |
| Comparing | Which Guildford builders are insured and give a workmanship warranty, and who quotes for free | Which Reading accountants already have clients in our sector, and what are their qualifications |
Your own name belongs in the coverage plan
Service questions tell you whether you get discovered. They do not tell you what the web has concluded about you, and that is a separate and highly concentrated signal. Somebody who has your card in their hand, or who was given your name by a friend, goes and asks about you directly. What comes back is the consensus, summarised.
- 1
Ask what you are known for
Tell me about [your business] in [your town]. Read the answer for what it gets wrong as much as what it gets right. Wrong opening hours, a service you stopped offering, a merged description of you and a similarly named firm: each is an entity problem with a specific fix. - 2
Ask whether you are any good
Is [your business] any good, what do people say about them? This one runs almost entirely on reviews and third-party mentions, so the answer is a direct readout of your reputation footprint rather than your website. - 3
Ask how you compare
How does [your business] compare with other [service] in [town]? The businesses named alongside you are the peer set the web believes you belong to, which is often not the peer set you would have chosen. See where those mentions live.
Two or three brand questions per location is plenty, and they are the cheapest diagnostic in the set. A thin, hedging or wrong answer to a direct question about your own name is the clearest possible evidence that the consensus problem is upstream of anything you write. It is also the fastest thing to fix, because the sources feeding it are ones you can go and correct today.
Breadth beats depth, and it is not close
Suppose you can afford forty checks a month. You can spend them as twenty keywords with two questions each, or as four keywords with ten. The second option feels more thorough and is close to worthless, because ten rephrasings of the same question mostly agree with each other. You will have bought ten slightly different views of one thing and no view at all of the sixteen services you did not check.
The variance that matters sits between keywords, not between phrasings of one. A business is rarely uniformly visible. It gets named consistently for the service it is known for, and is invisible for the two service lines it has been quietly trying to grow, which is exactly what you would predict from the way coverage in the training data and in retrievable sources tracks how well a thing is known.7 You only find that by asking about all of them.
Put the place in the query, or you are measuring someone else’s market
A real customer never has to say where they are. Their phone, their account and their history supply it, and Google resolves location at the moment of the search.10 Assistants carry their own version of that context, including what they remember about the person asking.11 When you send a question through a tracking tool, none of that exists, so if you do not write the place into the question, the engine answers for everywhere and nowhere. This is the single most common reason a tracking setup produces results that look inexplicably wrong.
City or town level is almost always right. Go finer, to a postcode or a street name, and the model has very little to work with and starts guessing. Go broader, to a county or a country, and you are measuring a national conversation dominated by national brands, which is not the market you compete in.
Ask every engine the same set
ChatGPT and Google’s AI Mode do not work the same way and will not give you the same answer. One turns your request into its own searches and assembles a reply from what comes back;6 the other fans your question out across subtopics inside Google’s own stack.5 A business can be well covered in one and thin in the other, and that gap is genuinely useful information about where your work is landing.
It is only useful if the questions are identical. Run one set on one engine and a different set on the other and you can never tell whether the difference in front of you is the engine or the question, which makes the comparison worse than not running it. Same ladder, every engine, every time.
One thing not to do: never read the order businesses appear in as a ranking. An answer engine is writing a sentence, not publishing a results page, and the name it happens to mention first is a feature of prose. There is no position to track. What you can track is how often you appear at all, and how much of the available mention space is yours against the competitors who keep coming up.
What to measure once the questions are running
Visibility rate
Out of every question run this period, how often you were named at all. The most honest headline number there is, and the one that moves when your work lands.
Share of voice
Of all the businesses named across your question set, how much of that space was yours. This is what replaces a rank position, and it is the number that survives a new competitor entering your market.
Sources cited
Which pages and directories the engine leaned on to build the answer. This is your to-do list: the places you need to be correct on, and often the reason a rival keeps getting named instead of you.
Read the trend rather than any single check, and understand why that is not a hedge. Identical requests to a model can return different text run to run, largely because the load on the server changes the batch a request is processed in, which changes the arithmetic slightly.12 That variation is in the infrastructure, not in your setup and not in your marketing. One check is an anecdote. Eight weeks of repeated checks is a measurement. Keep the responses word for word too, because a number without the reply behind it cannot be audited, and a number you cannot audit is one you should not put in front of a client.
Once a question is running, stop editing it
This one catches people out. Rewriting a question that has been running for two months does not improve your history, it ends it. Whatever trend line you had was keyed to that wording, and the new wording starts from nothing. The discipline is the same one that makes rank tracking meaningful: hold the ruler still, so the change you see is a change in the world rather than a change in the measurement.
So add rather than edit. New service, new keyword, new ladder. Want to test a different angle, add it as a rung and leave the existing ones alone. Review the whole set once a quarter, or whenever your service list genuinely changes, and expect most of it to stay exactly as it was.
How this is built into SearchOps
Everything above works with a spreadsheet and a lot of patience. We built it into the product because the patience is where it falls apart. Someone sets up a careful ladder in month one, adds four keywords in month three, never goes back to expand them, and the new services stay invisible in the reporting for a year before anybody notices.
In ChatGPT Visibility, the ladder is called Coverage, and it is stored against the location rather than rebuilt each time. You choose the depth once, ready to buy, comparing, or full funnel, and the questions are drafted from your tracked keywords using the angles that suit your sector, resolved from your business category with an override if we get it wrong. Because the plan is stored, adding keywords later is not a trap: next time coverage is read, new keywords are brought up to the depth you already chose, whichever route they arrived by. Brand evaluation is a switch on the same plan, so the questions about your own name run alongside the service ones. Our agent then improves the drafted wording using what it knows about your business, and the same set runs against every engine you have switched on, so the Gemini numbers and the ChatGPT numbers are actually comparable.
Two things it deliberately does not do. It does not claim to guess the exact sentence your customer will type, because that is not knowable and pretending otherwise would be selling you a story. And it will not quietly redraft questions that are already running, for the reason in the section above. What it does instead is make sure the space is covered, keep it covered as your keyword list grows, and show you the parts of it where the web has not yet made up its mind about you.
The audit, in one page
- Every question traces back to a keyword you already track or to something a customer actually said, in a call, a form, a chat or a review.
- The qualifiers your customers repeat, the deadline, the constraint, the worry, appear in the question set rather than only in your notes.
- Every question names a real place you trade in, qualified if the town name exists in more than one country.
- Every question is written the way a customer talks, demands included, rather than in keyword shorthand.
- Most questions sit at the point of decision. At most one per keyword is an early-journey question.
- Two or three questions ask about your business by name, so you can see the consensus directly.
- Breadth before depth: every service line that matters has at least two questions on it.
- The same question set runs against every engine you report on.
- You are reporting visibility rate, share of voice and cited sources, and never a position or a place in a list.
- Responses are kept in full, and the headline is the trend across repeated runs rather than the latest single check.
- Wording of running questions is frozen. New angles get added as new rungs, not as edits.
- Keywords added since setup have been expanded to the same depth as the originals.
None of this is exotic. It is the same discipline that made rank tracking useful twenty years ago: measure the things people actually do, hold the ruler still, and read the trend rather than the reading. What is new is that the thing people do is now typing a sentence rather than a fragment, that the engine turns that sentence into a search of its own, and that whether your name survives the process depends on what the rest of the web has been saying about you. Cover the space, and you can finally see which part of that is missing.
Keep reading
References
Every claim above that rests on someone else's work is marked inline and listed here, with a link to the original where we link out. Links last checked 30 Aug 2026.
- 1.Pew Research Center (June 2026). Americans and AI 2026: Chatbots, smart devices and views on impact. Survey of 5,119 US adults, 17 to 23 February 2026: about half now use AI chatbots, up from a third in 2024. Checked 30 Aug 2026.
- 2.Chatterji et al., National Bureau of Economic Research (September 2025). How People Use ChatGPT (Working Paper 34255). Privacy-preserving classification of ChatGPT conversations. Practical Guidance, Seeking Information and Writing account for close to 80% of them. Checked 30 Aug 2026.
- 3.Nielsen Norman Group (August 2006). Use old words when writing for findability. Search queries run to two or three words, and people reach for the familiar word rather than the marketing one. Checked 30 Aug 2026.
- 4.Zhao et al. (2024). WildChat: 1M ChatGPT interaction logs in the wild. A corpus of one million consented, real user conversations, collected precisely because researcher-written prompts do not resemble what people actually send. Checked 30 Aug 2026.
- 5.Google (March 2025). AI Mode in Search. Describes the query fan-out technique: one question becomes multiple related searches across subtopics and sources. Checked 30 Aug 2026.
- 6.OpenAI. Web search tool. How a model turns a request into its own searches and assembles an answer from what comes back. Checked 21 Aug 2026.
- 7.Kandpal et al. (2022). Large language models struggle to learn long-tail knowledge. A model's accuracy on a fact tracks how many documents supporting it appeared in pre-training. Thinly covered facts are thinly known. Checked 30 Aug 2026.
- 8.Lewis et al. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. Fetching documents at answer time reduces, without removing, a model's dependence on what it absorbed in training. Checked 30 Aug 2026.
- 9.Google. Tips to improve your local ranking on Google. Relevance, distance and prominence: the shared signals behind local results, whatever surface asks for them. Checked 21 Aug 2026.
- 10.Google. Understand and manage your location when you search on Google. Location is an input resolved at search time and changeable by the user, not a fixed property of the person asking. Checked 21 Aug 2026.
- 11.OpenAI. Memory FAQ. What ChatGPT carries between conversations, and what the person asking can switch off. Checked 21 Aug 2026.
- 12.He, Thinking Machines Lab (September 2025). Defeating nondeterminism in LLM inference. Identical requests differ run to run largely because server load changes the batch size a request is processed in, not because of anything you did. Checked 30 Aug 2026.