Answer the Business

Why the same question gets a different answer twice

AB
The Answer the Business team
· 8 min read

Your friend asked ChatGPT and you were there. You asked and you were not. Neither of you is wrong — and the disagreement is the most useful thing on the page.

Ask twice, get two answers

Somebody tells you they asked an assistant for the best in your trade in your town and you were named. You try it yourself an hour later and you are not there. One of you assumes the other is lying, or that something broke in between.

Nothing broke. Answer engines are not lookup tables, and the same question can produce a different answer minute to minute, run to run, phone to phone. Once you know why, two things follow: you stop panicking about any single answer, and you start asking much harder questions of anyone selling you a screenshot as evidence.

Reason one: the answer is generated, not retrieved

A search engine returns rows from an index. Ask twice, get the same rows.

A language model does something else. It produces the answer a word at a time, and at each step it is choosing from a distribution over the possible next words rather than reading a stored result. That choice involves sampling. Two runs of an identical prompt can diverge at the first word where two candidates were close, and once they diverge they stay diverged — a sentence that begins by describing what to look for in the trade ends up naming different businesses than one that begins with a list.

There is no setting an owner can reach that turns this off, and in the consumer apps there is no setting at all. Even where providers offer knobs that reduce variation, the answers are served from large distributed systems where identical inputs are not guaranteed to take an identical path. Treat run-to-run variation as a property of the medium, not a fault in it.

Reason two: the retrieval changed underneath it

For local questions, the model is not answering from memory. It runs a search first, reads what comes back, and writes from that.

So every source of variation in ordinary search is inherited whole. The index updates. A directory page is recrawled. A news item or a community thread rises for an hour and falls again. A review lands. The search step that fed one answer is simply not the same search step that fed the next one, and a business sitting just below the cut in what gets read is in some answers and out of others — with nothing about the business having changed at all.

This is also why being present in more of the sources engines lean on is more robust than being present in the best one. Breadth is what stops a single recrawled page deciding whether you exist that afternoon.

Reason three: the world actually changed

The least mysterious reason and the most common. Between the two runs, something real moved.

Your competitor collected four reviews. A directory finally corrected an address it had held wrong for two years. Somebody wrote a listicle for your town. Your own site changed. Answers about local businesses are assembled from sources that change every day, and they change asymmetrically — most weeks nothing about you moves while something about somebody else does.

And then there is who is asking

Two more variables sit outside the model entirely.

Location is the big one. Answer engines take the asker’s location seriously for local questions, and a question asked from the far side of a city is a different question. This is why an audit fixes the location deliberately rather than letting it float.

Phrasing is the other. The best in town, who should I call, who is open right now and who is cheap are four different questions, and they get four different answers. Somebody who tried one phrasing has one data point about one phrasing.

What a screenshot proves

Very little, and it cuts in both directions.

A screenshot of an assistant naming you is one sample of a distribution, taken at one moment, from one location, by one person whose account may carry memory of earlier conversations. It is pleasant and it is not evidence of a position. A screenshot of an assistant naming your rival instead is the same thing wearing a different mood.

Apply that to the marketing you are shown. A vendor demonstrating that they put a business into an answer, with a screenshot, has demonstrated that they ran the question once after the fact. The claim to ask about is not what one answer said; it is how many runs were taken, from where, with what question, and how many of them came back the same way. If the answer is a method, it can be checked. If the answer is an image, there is nothing there to check.

The same rule applies to us, which is why the report shows the answers themselves rather than the conclusions drawn from them.

What this means if you are measuring

It means a single answer is an anecdote. It might be a true anecdote and it is still not a measurement, because you cannot tell from one run whether you were absent or merely unlucky in that run.

Which puts a floor under any honest method: ask more than once, and report the rate rather than the verdict.

Every question in a full audit is asked of every engine three times, and the result recorded is how many of those runs named you. Named in three of three and named in one of three are different facts about your business. The first is a position. The second is a coin flip that happened to land your way, and an owner told only that they were present would be told something that will not hold next week. Both cells look identical in any tool that stores a tick.

Ranking is handled the same way, from the other end. Where you were named across several runs, the position kept is the best one you achieved — because being capable of the top of the list is the real finding, and averaging positions across runs where you did not appear at all would quietly punish you for variance rather than describe it.

Absence is recorded with the same care. A run where no business at all was named is not the same as a run where a rival was named and you were not — the first is a question the engines are answering vaguely, the second is a question you are losing to somebody specific. Flattened into one number they look alike, and the advice they deserve is not remotely the same.

What we hold still on purpose

Variance is worth measuring. Variance we introduce ourselves is worth eliminating, and the two get confused constantly.

Three things are pinned. The questions are generated from the category and the city alone, so the same market always produces a byte-identical set and this month’s audit is not measuring a different exam to last month’s. Answers are stored per question, per location, per engine, per run, per day, and never reused across days — a corpus that carried yesterday’s answers forward would freeze your score while the world kept moving, which is the flattering failure rather than the honest one. And the scoring itself is deterministic and pure: the same recorded answers produce the same score, today and in a year.

That last one is not fussiness. If a customer who changed nothing sees their number move because our arithmetic drifted, the number is finished — they will never again be able to tell our noise from their progress.

The same discipline is why every audit carries the version of the scoring engine that produced it, and why scores from different versions are never compared. When what we measure changes, the version changes with it, so a jump you did not earn cannot be reported to you as progress you did.

So your score moved and you did nothing

Some of that movement is real and some is the medium. In order, here is what to check.

First, whether the two scores are even comparable — different scoring versions measure different things and are not charted together for exactly this reason.

Second, the cells rather than the total. A score that fell two points because one medium-band question flipped from named-in-two to named-in-one is telling you something very different from a score that fell two points because a wrong closing time started being repeated. The first is variance you should not act on. The second is a job for this afternoon.

Third, the pillars that cannot flicker. Structured data, an llms.txt file, crawler access, whether your site states prices and a service area — none of these change unless somebody changed them. If those moved, somebody redeployed a site or edited a profile, and that is the most findable kind of cause.

Fourth, whether the market moved rather than you. A score is a position among other businesses, and a rival who collected a season of reviews or finally put prices on their site can push you down a list you did nothing to fall on. That is a real finding, not a glitch, and the answer to it is work rather than reassurance.

What to do about a single bad answer

Very little, on its own.

Chasing one answer is how people end up optimising for noise: rewriting a homepage because one run of one engine on one afternoon named three rivals. The next run may well have named you, and now the site has been changed for a reason that was never real.

What is worth acting on is a pattern that survives repetition. Absent from every run of a high-value question across several engines is a finding. Named in a third of the runs is a different finding — you are on the edge of what the engines consider a good answer, and the work is to make the sources agree about you until you are comfortably inside it. Both of those are things you can only see by asking repeatedly, in a fixed location, with questions that do not change between runs.

That is the whole reason the audit is built the way it is, and the method is written down so you can check that it does what it says. What comes out the other end is a score you can take apart, which is the only kind worth having when the underlying thing being measured refuses to sit still.

See where you stand right now.

A free scan across every engine we probe. No card, no login.

Scan my business — free