Answer the Business

How we ask: the method behind every audit

AB
The Answer the Business team
· 8 min read

A visibility score is only as good as the questions behind it. Here are ours in full — including the rules that cost us points we could have quietly claimed.

The questions are the measurement

Everything downstream of the question set is arithmetic. Ask the wrong things and you will produce a confident number about nothing — a business named in questions nobody types is not visible in any way that pays, and a business absent from questions everybody types has a problem no average will show you.

So the questions come first, and they are generated under three rules that are load-bearing rather than stylistic. Each one costs something. That is usually how you tell a rule from a preference.

Rule one: never name the business

Not once, anywhere in the run.

The obvious way to test whether an engine knows about you is to ask it about you. It is also the fastest way to manufacture the answer you were hoping for. Name a business in the prompt and you have handed the engine the retrieval term, the subject and the sentiment; it will find something to say, and what it says tells you almost nothing about whether it would have reached for you unprompted.

Every question is therefore phrased the way a stranger who has never heard of you would ask it. Someone who needs the job done, knows the trade and the town, and knows nothing else. If you are not in that answer, you are not in the answer a real customer gets either — which is the only finding worth paying for.

Rule two: everyone in a market gets the same questions

Questions are generated from the category and the city, and nothing else. No business name, no owner detail, no site content. Two businesses in the same trade in the same town produce a question set that is identical down to the byte.

That has one obvious benefit and one that pays for the product. The obvious one is comparability: when your neighbour scores higher, it is not because they were asked easier questions. The second is that an engine’s answer to a question about the best in town is not about you specifically, so it can be asked once and read for everyone in that market. A question that varied per business would silently cost a full-price probe every time, and the report would cost what the probes cost.

Rule three: every category gets the full set

The category is whatever Google says the business is, which means the library has to cover every category Google knows — including the ones we have never seen.

That guarantee used to be false and shipping. Seven recognised categories got the full depth; everything else got the general questions only, which ran out well short of the depth the pricing page was advertising. A florist was sold one number of questions and received another. Two smaller defects sat underneath it: a business Google had no category for was quietly probed as though it were a plumber, and Google’s own wording never matched our internal keys, so the categories most likely to be misnamed got the thinnest treatment.

The library now has a general set large enough to fill a full audit on its own, for any category on earth, plus intent questions for several dozen recognised trades layered on top. The general set is the floor; the trade-specific questions are an upgrade, never the coverage. A category we do not recognise gets a valid categorical measurement rather than somebody else’s questions, and a test fails the build if any category cannot reach full depth.

Where the category comes from, and how to overrule it

It comes from Google’s own primary type, because a list we maintain by hand would be wrong about somebody within a week.

Google is also confidently wrong often enough to end the conversation. A photocopy shop that Google had filed as a manufacturer was audited as one — seven engines asked to name the best manufacturer in the city, and a result that looked broken to the one person who could see why it was not.

So the owner can correct it, the correction is recorded as having come from a human, and the next scan cannot quietly restore Google’s answer. The correction is folded into the shared market key so a retyped category cannot fork a market on capitalisation alone, and it is pinned to the audit at planning time so a change mid-run cannot rewrite a measurement already counted. Scores are only ever compared within the same category, for the same reason they are only compared within the same scoring version.

Buying questions and research questions

Two kinds of question, both probed identically, tagged so the report can tell them apart.

A buying question is somebody ready to choose: who is the best in town, who should I call, who is open right now. Being named there is the sale, and being absent is a lost one.

A research question is the same person twenty minutes earlier, working out what the job involves and what it should cost. Nobody is choosing yet — but this is where an engine reaches for a page that explains something, and the business with a real pricing or guidance page is the one it reaches for. It is also the only lane an owner can win by writing something, which makes it the lane the fix list can actually act on.

A fixed share of every full audit is reserved for research questions, taken from inside the total rather than added on top. Buying coverage and research coverage compete for the same budget, which is honest about the trade-off rather than pretending there isn’t one.

Not every question counts the same

Each question is banded by roughly how often it gets asked — high, medium or low — and that band becomes its weight in the presence pillar. A mention in a high-band question is worth about three times a mention in a low-band one.

The high band is kept deliberately small. Most real demand concentrates into a handful of phrasings, and spreading the weight evenly across every variation would flatten the one thing that makes the presence number mean anything.

The questions are also ordered the way a buyer actually moves: broad discovery first, then urgency, then qualification, then price. That ordering shows up in your report as the shape of your problem — plenty of businesses are visible in discovery and vanish the moment the question mentions money.

What we ask, and what we ask it with

We probe each provider’s own API with its native grounded search switched on, and with the location set to the market being measured. Location is the whole ballgame for local intent; without it the model answers for nowhere in particular and the result is noise.

A full audit puts twenty-four questions through five engines three times each — three hundred and sixty answers, recorded verbatim. The free scan is a much smaller version of the same thing: a handful of questions, once each, across fewer engines.

Here is the disclosure that most of this industry leaves out. A grounded API response is a proxy for what a person sees in the consumer app. It is the most reproducible measurement available and it is close to what a customer gets, but it is not identical — the apps personalise, carry memory, and ship features unevenly. Anyone claiming to measure the consumer app itself is measuring something they cannot observe. An engine we did not probe is never drawn as a miss; it is simply not drawn.

Asking more than once

Every question is asked of every engine more than once, and the score records how many of those runs named you rather than whether any of them did.

That is not padding. Ask an engine the same question twice and you will often get two different answers, for reasons that have nothing to do with your business — which is worth understanding before you read your own report. One answer is an anecdote. A rate across repeated runs is a measurement.

Answers are stored per question, per location, per engine, per run, per day, and never reused across days. Engines change their minds; a corpus that carried yesterday’s answers forward would quietly freeze your score while the world moved.

Reading the answer: two readers, one verdict

Deciding whether you were named sounds trivial and is not. Business names are short, generic and frequently shared with a street, a suburb or a rival two towns over.

So each answer is read twice. A deterministic string match — case folded, accents stripped, punctuation normalised, link URLs discarded so a bare directory link does not count as a mention — finds the name literally in the prose and, where the answer is a ranked list, the position it appears at. Separately, a model reads the same answer and gives its own verdict, the businesses named, the sentiment and the specific claims made about you.

You count as named only when both agree. Where they disagree, the disagreement is recorded rather than resolved by whichever reader is more convenient. A model asked whether a business was mentioned will happily say yes about a paraphrase, a near-miss name, or nothing at all. The string match cannot, which is exactly what makes it useful as a check on the part of the pipeline that cannot be proven.

Wrong facts: the bias is toward silence

Claims an engine makes about you — hours, contact details, location — are checked against what Google actually holds, and a claim is only ever called wrong when it asserts a value we hold and that value confidently disagrees.

Everything vague, partial or ungrounded is skipped. Skipped is not the same as correct: it is outside what we checked, the same discipline we apply to engines we did not ask.

What this method cannot see

It cannot see what any individual customer was shown, because that answer was personalised and is gone. It cannot see the consumer apps directly. It cannot tell you your share of a market, because nobody outside the providers can. And it cannot promise that fixing what it finds will put you in the answer, only that the things it found are the things standing between you and it.

What it can do is show you the questions, the answers they produced, and exactly how those became a number.

See where you stand right now.

A free scan across every engine we probe. No card, no login.

Scan my business — free