How We Measure AI Visibility — the Full Method, Published
By Udaay Sikder
Most agencies selling "AI search optimization" will show you a proprietary score. We won't. A score you can't reproduce is an invoice justification, not a measurement. Here is our entire method — precisely enough that you can run it yourself, on your own business, today, without us.
The problem with measuring AI answers
Ask ChatGPT the same question twice and you can get different answers. That's not a flaw to hide; it's the central fact any honest measurement must be built around. AI assistants are non-deterministic — so a single query proves almost nothing, a screenshot of one good answer proves even less, and anyone selling you a "#1 ranking in ChatGPT" is selling a coin flip they photographed mid-air.
The honest unit of measurement is a rate: across a fixed set of realistic questions, asked repeatedly, under controlled conditions — how often are you named?
The method, step by step
1. A frozen prompt set. Before measuring anything, we write ~20 prompts in the voice of your actual buyers — not keyword language. Real people don't ask "best HVAC Indianapolis"; they ask "my AC just died and it's summer, who should I call in Indianapolis?" The set mixes direct requests, need-based scenarios, trust-modified asks ("most reliable…"), suburb-specific variants, and urgency cases.
Frozen means frozen: the set is fixed on day one, shared with the client in full, and never changes without written sign-off. If the yardstick can quietly change, the measurement is theater.
2. Three engines. Each prompt runs on ChatGPT, Gemini, and Perplexity — the assistants consumers actually use for recommendations. Same prompts, all three.
3. Three runs each. Every prompt runs three times per engine, each in a fresh session with no history. Because outputs vary, one run is an anecdote; three begin to be a rate. (More runs would be better statistics; three is the honest floor that keeps monthly measurement affordable — and we say so rather than pretending three is thirty.)
4. Strict run rules. Fresh session every time. Prompts pasted verbatim, character for character. No follow-ups, no clarifications, no "are you sure?" Refusals and generic answers are recorded as data (none named), not discarded. Every run is screenshotted and archived, so any number in any report can be traced to its evidence.
5. What gets recorded. For every run: which businesses were named, in what order, and — where the engine shows them — which sources it cited. The citations matter as much as the mentions: they reveal where an engine gets its local answers (review platforms, directories, local press), which is what turns measurement into a work plan.
6. What gets reported. Three numbers, monthly, on the same frozen set:
- Mention rate — the share of runs in which you were named at all
- Position profile — when named, where in the answer
- Citation sources — which platforms the engines credited
Baseline first, then the same measurement every month. Movement on a frozen yardstick is progress; anything else is noise or narrative.
The limitations, stated plainly
Every measurement method has them; here are ours. Personalization: consumer AI apps adapt to logged-in users, so your own results may differ from a clean session's — we measure clean sessions because that's the shared reality of new customers. Temporal drift: engines update constantly; a rate is a snapshot, which is exactly why the set re-runs monthly instead of being sold as a permanent fact. Coverage: three engines is not all engines; it's the three with meaningful consumer recommendation use today, revisited as the market moves. And the one that governs everything: no one can guarantee AI results. The systems are non-deterministic and externally controlled. What can be honestly sold is the measurement, the foundational work the engines demonstrably reward — crawlable pages, consistent business data, real reviews, genuine content — and the discipline of re-measuring on a yardstick that doesn't move.
Run it yourself
You don't need us to start: write five questions your customers would actually ask, run each once in ChatGPT and Gemini in fresh sessions, and count who gets named. That ten-minute version is coarse, but it will tell you whether you have a visibility problem — free. (Our five-minute walkthrough shows the exact steps, and the technical half of the check — whether AI crawlers can even read your site — runs free on Crawlmouse.)
When you want the full instrument — the frozen set, three engines, three runs, monthly, with the evidence archived — that's the service, and now you know exactly what you'd be buying, because you've read the whole method.
This methodology is applied in our sample engagements and in every client report we produce. Questions about it — including the parts you think are wrong — are welcome: talk to us. A method that can't take criticism isn't one.
Related reading
- The Four-Vector Screen: Where the Money Hides in a Custom Injection Molder
The Four-Vector Screen: Where the Money Hides in a Custom Injection Molder
A sample engagement for a fictional Indiana molder: four money vectors screened with labeled numbers and ranked, including the row where the project fails.
September 15, 2026
- The Quote That Took Six Days: AI Quoting for an Indiana Machine Shop
The Quote That Took Six Days: AI Quoting for an Indiana Machine Shop
A sample AI engagement for a fictional Kokomo job shop: the quoting bottleneck measured, a labeled ROI model with its negative row, an AI-assisted estimating workflow, and payback inside five months.
September 11, 2026
- The Proposal That Ate March: AI for a Custom Machine Builder's Quoting Bottleneck
The Proposal That Ate March: AI for a Custom Machine Builder's Quoting Bottleneck
A sample AI engagement for a fictional Ontario automation integrator: the unpaid engineering inside every proposal, a labeled ROI model with its negative row, and payback inside seven months.
September 11, 2026
Nahl Technologies