What an AI visibility audit cannot tell you
Every tool publishes what it measures. Almost none publishes what it misses — which is the half that decides whether the first half means anything.
The half nobody publishes
Every measurement tool tells you what it measures. Almost none tells you what it misses, which is unfortunate, because the second list is what decides whether the first one means anything.
So here is ours. Not caveats buried in a methodology page — the actual list of things a visibility audit cannot do, including the ones a customer would be within their rights to have assumed it could.
If any of these turn out to be the thing you needed, you now know before you pay rather than after.
It cannot tell you what your customer was shown
This is the big one, and it is structural rather than a gap we are working on.
The audit probes each provider’s own interface with grounded search switched on and the location set to your market. That is the most reproducible measurement available, and it is close to what a person gets. It is not identical. The consumer apps personalise, carry memory of earlier conversations, and roll out features unevenly across accounts and countries.
So when a customer of yours asked an assistant last Tuesday and got a particular answer, that answer is gone. Nobody can retrieve it — not us, not any competitor, not the provider’s support desk. Anyone claiming to measure the consumer app itself is describing something they cannot observe, and it is worth asking them how.
What the audit gives you instead is a controlled, repeatable version of that question, asked the same way every time, which is the only thing that can be compared month over month. A controlled measurement is not the same as the real event. It is what you use when the real event is unobservable, and the honest framing is a proxy rather than a recording.
Location granularity is part of the same limit. The audit asks from your market, because a market is the unit a business competes in and a fixed location is what makes two runs comparable. Your customers ask from a particular street, on a particular network, and an engine takes that seriously. Being named across a city does not guarantee being named from every corner of it, and no audit priced for a local business can probe every corner.
It cannot tell you whether an answer produced a call
There is no referrer on a spoken recommendation.
When somebody reads a paragraph naming you, taps the number and calls, nothing in your analytics records that the sentence came from an assistant at all — let alone which one. The most valuable outcome of being named is also the least traceable, and no tool on the market has solved this, whatever the dashboard implies.
This has a practical consequence worth stating plainly: nobody can currently show you a return on this work as a revenue figure. What can be shown is whether you are named, how often, in which questions, against which rivals. Treating that as a leading indicator is reasonable. Treating it as proof of income would require a link that does not exist yet.
It cannot tell you your share of the market
We do not publish a share-of-answers number, because nobody outside the providers can measure one and the ones in circulation are estimates dressed as data.
That refusal has a cost. It is a chart customers ask for, and the pages that do show one look more authoritative than this one. It stays refused because the alternative is manufacturing a number, and a tool that manufactures one number should not be believed about the rest.
It cannot promise that a fix puts you in the answer
Each fix in the report carries an impact figure, and the figure is honest about what it is: the audit re-scored with that one gap closed, and the difference between the two scores. It is not estimated by hand and it is not a guess.
It is also not a forecast about the engines. Closing a gap changes what the score measures about you; it does not oblige an assistant to name you next week. The gap between those two is real and nobody controls it — publishing structured data does not compel a model to reach for your site, and no honest vendor can promise placement, ourselves included. We say so where a buyer will read it rather than in the terms.
What the impact number does tell you is where the cheapest points are, and the fixes are ranked by impact against effort so the ordering means something even if the absolute figure moves.
It cannot tell you about facts it did not check
Wrong-fact detection only fires when an engine asserts a value we hold and that value confidently disagrees. Hours, contact details, location. Anything vague, partial or ungrounded is skipped.
Skipped is not correct. A claim we could not verify sits outside what we checked, and the report says so rather than counting it as clean. The bias runs toward silence deliberately: telling you an engine got your phone number wrong when it did not is worse than missing one it did, because the first mistake costs you your belief in every other line on the page.
The same discipline applies to engines. One we did not probe is never drawn as a miss — it is simply not drawn, and the scoring weight redistributes across the engines that did run so you are not marked down for a question nobody asked.
It cannot score everything it can see
Some things are measured, shown to you, and deliberately given no points.
Your photo count is the clearest example. The data source caps the number of images it returns, so what comes back is a floor rather than a count — and a pillar built on a number we cannot swear is a count is exactly what our own rule against unmeasured figures forbids. It ships as advice with no points attached, alongside the profile fields the source will not report at all.
That is the general shape of it: where we can see something but cannot count it honestly, it becomes advice rather than arithmetic.
It cannot compare you to a business in a different category
Scores are only compared within the same category and the same city, because the question set is generated from those two things. A dentist and a roofer were asked different questions in different markets, and comparing the results would be comparing two different exams.
Nor are scores compared across scoring versions. When what the audit measures changes, the version changes with it, and a business that did nothing would otherwise appear to have gained or lost points it never earned.
It cannot tell you what a single answer means
An answer is one sample. Every question is asked several times precisely because the same question returns different answers, and a single run cannot distinguish being absent from being unlucky.
So the report speaks in rates rather than verdicts, and the honest reading of one cell is a probability rather than a fact about your position. Anyone who shows you a single screenshot as evidence — for you or against you — is showing you one draw from a distribution.
It cannot tell you why an engine chose somebody else
It can tell you that it did, which rival it named, and which sources it leaned on while doing it. That is a lot, and it is not the same as a reason.
Nothing observable from outside a model explains why one business made the paragraph and another did not. What comes back is an answer and a set of citations, and any account of the reasoning is a reconstruction — a plausible one, built from what the sources say, but a reconstruction. When a report tells you a rival is winning a question because their site states prices and yours does not, that is an inference from a pattern rather than a confession from the engine.
The honest version is to keep the two separated on the page: here is what was observed, here is what we think explains it, and the second one is argued rather than measured. A tool that presents the inference in the same typeface as the observation is quietly upgrading its own guesses.
It cannot do the work
The audit is a diagnosis. Every fix it names has to be carried out by somebody with the passwords.
We cannot edit your Google profile, claim a directory listing on your behalf, persuade an aggregator to correct a record it has held wrong for two years, or make anybody mention you in a community thread. Some of those are slow. A few are genuinely outside anyone’s control, which is why the report separates what you can close this week from what will take a season.
That boundary is deliberate rather than a limitation we intend to remove. The moment a tool both scores you and edits the sources it scores you on, its numbers stop being independent of its own work — and you would have no way left to check either.
Why publish this at all
Because the alternative is worse, and because the list is short enough to survive being read.
Being believable about what we measure is the only durable advantage a product like this has. Every competitor can copy a feature. None of them can copy a reputation for refusing to overstate, and the way you build one is by writing the limits down before a customer discovers them on their own.
There is also a selfish version of the argument. A customer who knows what the audit cannot do will not go looking for it in the report, will not misread a number as a promise, and will not be disappointed by a gap they were never told about. The support burden of an honest product is lower than the support burden of an impressive one.
If you want the other half — the questions, the engines, the readers, and how a mention becomes a number — it is in the method and in what the score is made of. Read them together. Either one alone tells you less than half the story.
A free scan across every engine we probe. No card, no login.
Scan my business — free