How We Measure AI Visibility — the Full Method, Published
By Udaay Sikder
Most agencies selling "AI search optimization" will show you a proprietary score. We won't. A score you can't reproduce is an invoice justification, not a measurement. Here is our entire method — precisely enough that you can run it yourself, on your own business, today, without us.
The problem with measuring AI answers
Ask ChatGPT the same question twice and you can get different answers. That's not a flaw to hide; it's the central fact any honest measurement must be built around. AI assistants are non-deterministic — so a single query proves almost nothing, a screenshot of one good answer proves even less, and anyone selling you a "#1 ranking in ChatGPT" is selling a coin flip they photographed mid-air.
The honest unit of measurement is a rate: across a fixed set of realistic questions, asked repeatedly, under controlled conditions — how often are you named?
The method, step by step
1. A frozen prompt set. Before measuring anything, we write ~20 prompts in the voice of your actual buyers — not keyword language. Real people don't ask "best HVAC Indianapolis"; they ask "my AC just died and it's summer, who should I call in Indianapolis?" The set mixes direct requests, need-based scenarios, trust-modified asks ("most reliable…"), suburb-specific variants, and urgency cases.
Frozen means frozen: the set is fixed on day one, shared with the client in full, and never changes without written sign-off. If the yardstick can quietly change, the measurement is theater.
2. Three engines. Each prompt runs on ChatGPT, Gemini, and Perplexity — the assistants consumers actually use for recommendations. Same prompts, all three.
3. Three runs each. Every prompt runs three times per engine, each in a fresh session with no history. Because outputs vary, one run is an anecdote; three begin to be a rate. (More runs would be better statistics; three is the honest floor that keeps monthly measurement affordable — and we say so rather than pretending three is thirty.)
4. Strict run rules. Fresh session every time. Prompts pasted verbatim, character for character. No follow-ups, no clarifications, no "are you sure?" Refusals and generic answers are recorded as data (none named), not discarded. Every run is screenshotted and archived, so any number in any report can be traced to its evidence.
5. What gets recorded. For every run: which businesses were named, in what order, and — where the engine shows them — which sources it cited. The citations matter as much as the mentions: they reveal where an engine gets its local answers (review platforms, directories, local press), which is what turns measurement into a work plan.
6. What gets reported. Three numbers, monthly, on the same frozen set:
- Mention rate — the share of runs in which you were named at all
- Position profile — when named, where in the answer
- Citation sources — which platforms the engines credited
Baseline first, then the same measurement every month. Movement on a frozen yardstick is progress; anything else is noise or narrative.
The limitations, stated plainly
Every measurement method has them; here are ours. Personalization: consumer AI apps adapt to logged-in users, so your own results may differ from a clean session's — we measure clean sessions because that's the shared reality of new customers. Temporal drift: engines update constantly; a rate is a snapshot, which is exactly why the set re-runs monthly instead of being sold as a permanent fact. Coverage: three engines is not all engines; it's the three with meaningful consumer recommendation use today, revisited as the market moves. And the one that governs everything: no one can guarantee AI results. The systems are non-deterministic and externally controlled. What can be honestly sold is the measurement, the foundational work the engines demonstrably reward — crawlable pages, consistent business data, real reviews, genuine content — and the discipline of re-measuring on a yardstick that doesn't move.
Run it yourself
You don't need us to start: write five questions your customers would actually ask, run each once in ChatGPT and Gemini in fresh sessions, and count who gets named. That ten-minute version is coarse, but it will tell you whether you have a visibility problem — free. (Our five-minute walkthrough shows the exact steps, and the technical half of the check — whether AI crawlers can even read your site — runs free on Crawlmouse.)
When you want the full instrument — the frozen set, three engines, three runs, monthly, with the evidence archived — that's the service, and now you know exactly what you'd be buying, because you've read the whole method.
This methodology is applied in our sample engagements and in every client report we produce. Questions about it — including the parts you think are wrong — are welcome: talk to us. A method that can't take criticism isn't one.
Related reading
- We Audited 187 Small-Business Websites. Half Have Pages Nothing Links To.
We Audited 187 Small-Business Websites. Half Have Pages Nothing Links To.
Original data from 187 real small-business websites: 50% have orphan pages, 75% over-optimize anchor text, and the smallest sites have the weakest structure. Small-business website statistics the big SEO studies don't cover — with method and limitations published.
August 16, 2026
- The $118,000 Voicemail: AI After-Hours Intake for an Indianapolis HVAC Company
The $118,000 Voicemail: AI After-Hours Intake for an Indianapolis HVAC Company
A sample AI engagement for a fictional Indianapolis HVAC contractor: phone-log analysis of a 31% missed-call rate, the labeled ROI model with its negative row, an intake system architecture, and a sub-nine-month payback.
August 13, 2026
- Drowning in Check Calls: AI Document Automation for an Indianapolis Freight Brokerage
Drowning in Check Calls: AI Document Automation for an Indianapolis Freight Brokerage
A sample AI engagement for a fictional Indianapolis freight brokerage: an inbox where 72% of email is mechanical, the labeled automation model with its negative row, a three-part document system, and the capacity case that beats the labor case.
August 13, 2026
Nahl Technologies