A plastic surgery practice measures AI visibility by running a fixed prompt set through each assistant on a schedule, three runs per prompt per engine, and scoring two numbers: coverage, the share of runs naming the practice at all, and win rate, the share where the practice ranks first. Every run is logged, including the runs that name nobody.
The ClinicAds pillar post on how plastic surgeons get cited by ChatGPT, Perplexity, and Google AI Overviews covers what to change on a practice site and off it. This post answers the question that pillar raises and leaves open: how a surgeon knows whether any of it worked. One sentence of setup on mechanics, then the measurement. An assistant assembles each answer from sources retrieved at query time, which is why the identical question asked twice can return two different shortlists, and why an AI visibility program has to be built on repetition rather than on a screenshot.
- AI visibility is measured with two numbers, not one. Coverage is the share of logged runs where the practice is named at all. Win rate is the share where the practice is the first or sole recommendation.
- A single run is a sample, not a verdict. On 2026-08-13 ClinicAds ran the same surgeon-selection query through three ChatGPT configurations and received three different agency shortlists, and a first-place result confirmed two days earlier did not reproduce in any of them.
- A usable prompt set for a plastic surgery practice is 12 to 20 selection-intent queries across four axes, run three times per query per engine, with every run logged whether or not the practice is named.
- Logging only the favorable runs produces a number that cannot be defended. ClinicAds holds its own public claim to the weaker wording until a four-part gate passes: 30 or more logged runs, 2 or more engines, a 21-day span, and the practice named in 70 percent of runs.
- A site-side AEO audit score measures whether a practice is optimized. It does not measure whether the practice is winning. The two numbers move independently and should never be reported as one.
How do you measure AI visibility for a plastic surgery practice?
AI visibility is measured by fixing a prompt set, running it on a schedule, and logging every run. A practice picks 12 to 20 selection-intent queries, runs each one three times per engine in a clean session, records whether the practice was named and in what position, and computes coverage and win rate from the log. The prompt wording stays frozen between cycles, because changing the wording changes the measurement.
The protocol ClinicAds uses on its own category is documented in the aeo/ROUTINE.md file in this repository and is deliberately unglamorous. Incognito or temporary-chat session, location set and recorded, three runs per query per engine, every run written to the log whether or not ClinicAds appears. That last rule is the one that gives the number its value.
- Step 1: write 12 to 20 selection-intent prompts and freeze the wording
- Step 2: pick 2 or more engines, at minimum ChatGPT and Perplexity
- Step 3: run each prompt 3 times per engine in a clean, logged-out session
- Step 4: log every run with date, engine, configuration, position, and the competitors named
- Step 5: compute coverage and win rate across the full matrix, then repeat weekly
What questions belong in the prompt set?
The prompt set should hold selection-intent queries, meaning questions where a patient is choosing a surgeon rather than researching a procedure. A query like how long is deep plane facelift recovery is answered in the pane with no practice named, so it measures nothing about a practice. A query like who are the best board-certified plastic surgeons in Sacramento for facelifts returns names, and names are what can be scored.
Four axes cover the realistic ways a patient phrases a selection question. A prompt set built on one axis produces a flattering number, because most practices are visible on their own brand name and invisible everywhere else.
- Include at least 2 prompts naming a direct local competitor; comparison queries are often the least contested
- Exclude pure procedure-research queries; they answer in-pane and name no practice
- Record the location the session reports, since a local shortlist changes with it
| Axis | What it tests | Example phrasing | Prompts |
|---|---|---|---|
| Category and geography | Whether the practice appears in an unaided local shortlist | Best plastic surgeons in [metro] for [procedure] | 5-6 |
| Credential and specificity | Whether the practice surfaces on the qualifier a patient screens with | Board-certified [procedure] surgeon in [metro] who does revisions | 4-5 |
| Problem language | Whether the practice is retrieved on the patient's own words rather than clinical terms | I had a nose job I am unhappy with, who should I see | 3-4 |
| Comparison and alternatives | Whether the practice appears when a named competitor is the starting point | [Competitor practice] alternatives in [metro] | 2-3 |
What is the difference between coverage and win rate?
Coverage is the share of logged runs where the practice is named anywhere in the answer. Win rate is the share where the practice is the first or the sole recommendation. A practice can hold 60 percent coverage and a 0 percent win rate, which means assistants know the practice exists but never lead with it. Those two conditions have different fixes, so collapsing them into one visibility number hides the problem.
Coverage responds to being present on the third-party surfaces an assistant retrieves. Win rate responds to entity clarity and to how decisively the sources describe what the practice is known for. A surgeon named fourth in a list of six is not close to winning; a surgeon named fourth is competing against whichever source the assistant leaned on for the ordering.
| Coverage | Win rate | What it means | Where the work goes |
|---|---|---|---|
| 0% | 0% | The practice is absent from the retrieved source set entirely | Third-party surfaces: directories, society listings, review platforms, local press |
| 20-50% | 0% | Named on some phrasings, never led with | Entity consistency and a sharper on-page definition of what the practice is known for |
| 60%+ | 10-25% | Established in the source set, competing on ordering | Depth on the specific procedure or qualifier being lost, plus recency of cited sources |
| 60%+ | 40%+ | Category position is real | Hold it: re-run weekly and watch for competitor movement |
Why is one run not enough to settle it?
One run is a sample because assistants do not return a stable answer to a repeated question. ClinicAds measured this on its own agency category on 2026-08-13. A prospect had confirmed a first-place ChatGPT result for a specific query on or before 2026-08-11. Two days later that exact query, re-run in three configurations, did not name ClinicAds in any of them, and each configuration returned a different set of competitors.
The capture below is a method demonstration on the agency category, not plastic surgery practice data, and it is one dated observation rather than a trend. It is included because the variance itself is the finding a practice needs before it starts screenshotting single answers.
- Across the full 12-cell matrix captured that day, coverage was 0 of 12 and win rate was 0
- The three ChatGPT captures shared no common agency in their top position
- A result confirmed 2 days earlier failed to reproduce in every replay
| Capture | Engine and configuration | ClinicAds named | Top of the shortlist returned |
|---|---|---|---|
| Prospect session, on or before 2026-08-11 | ChatGPT, logged out, default mode | Yes, ranked first | ClinicAds, then five other agencies |
| Replay 1, 2026-08-13 | ChatGPT, temporary chat, extended reasoning | No | Rosemont Media, then three agencies |
| Replay 2, 2026-08-13 | ChatGPT, temporary chat, instant mode | No | Etna Interactive, then four agencies, none shared with Replay 1 |
| Replay 3, 2026-08-13 | Google classic results, no AI Overview triggered | No | Five third-party listicles, no practice-level answer |
How often should a practice re-run the check?
Weekly for the priority queries and monthly for the full matrix. Weekly is frequent enough to catch a competitor entering the retrieved source set and infrequent enough that a practice is not spending an afternoon a day on it. A monthly full-matrix run recomputes coverage and win rate across every prompt and every engine, which is the number that belongs in a quarterly review.
Engine ranking checks stay manual. ChatGPT, Perplexity, and Google AI Overviews all restrict anonymous automation, and an automated check that returns degraded or challenged results produces a number that looks like visibility and measures nothing. ClinicAds automates the site-side audit daily and keeps the engine checks human-run for that reason.
- Weekly: 4 to 6 priority prompts, 3 runs each, 2 engines, roughly 30 logged runs
- Monthly: the full 12 to 20 prompt matrix, recompute coverage and win rate
- Daily: the site-side audit only, which checks schema, crawl access, and answer-block integrity
- After any site change to a practice bio or service page, re-run the affected prompts inside 14 days
What makes an AI visibility number untrustworthy?
Four failure modes account for most AI visibility numbers that do not survive scrutiny. The first is selective logging. A practice that records the runs where it appears and discards the rest reports a coverage figure that is arithmetically meaningless, and ClinicAds treats that as worse than reporting no number at all. The second is measuring in a logged-in, personalized session, where prior conversation history influences the answer and the result describes one account rather than the category.
The third is quietly rewording a prompt between cycles, which resets the baseline while appearing to continue it. The fourth is reporting a site-side audit score as if it were a ranking. A daily audit score covers schema health, vocabulary coverage, crawl surface, and extractability, all verifiable from the site itself. A practice can hold a 95 out of 100 site score and 0 percent coverage on the same day, because one measures readiness and the other measures outcome.
ClinicAds applies the same standard to its own public claims. The site says named by AI assistants rather than consistently named until a four-part gate passes: 30 or more logged runs, 2 or more engines, a 21-day span, and the practice named in 70 percent of runs. As of 2026-08-28 that gate has not passed, so the weaker wording stands.
- Log every run, including the ones naming no practice at all
- Use incognito or temporary chat, and record the location the session reports
- Freeze prompt wording; a reworded prompt starts a new baseline
- Report the site audit score and the coverage score as two separate numbers
What is this worth to a practice in booked cases?
The measurement is worth what the visibility is worth, which for a plastic surgery practice is measured in booked consultations rather than in citations. ClinicAds works to a planning range of $80 to $150 per booked consult on managed paid media at $5,000 to $10,000 per month, at 5 to 10x return on ad spend. Those are agency averages, not guarantees. An assistant recommendation arrives with none of that media cost attached, which is why the surface is worth tracking even while it sends a fraction of the volume paid search does.
Coverage moves over quarters rather than weeks, because it depends on third-party sources being published, indexed, and then retrieved. The reason to start logging now is that a practice cannot tell whether a change worked without a dated baseline to compare against, and the baseline is the cheapest part of the program.
What is a good AI visibility score for a plastic surgery practice?
There is no published benchmark, because the metric depends entirely on the prompt set a practice chose. The usable comparison is against a practice's own dated baseline and against the competitors appearing in the same logged runs. A practice named in 60 percent or more of runs across two engines, with a win rate above 25 percent, holds a real position in its metro.
Can AI visibility tracking be automated?
The site-side portion can. Schema validity, crawl access for AI user agents, llms.txt presence, and answer-block integrity are all verifiable from the site and ClinicAds audits those daily. The engine portion cannot be reliably automated, because ChatGPT, Perplexity, and Google restrict anonymous automated querying and return challenged or degraded results, so those runs stay human-run.
How many runs are needed before the number means anything?
At least 30 logged runs across 2 or more engines spanning 21 days or more. Below that, run-to-run variance dominates. ClinicAds saw a first-place result on one query fail to reproduce across three configurations two days later on 2026-08-13, which is the practical argument for the sample size.
Does a high AEO audit score mean a practice is being cited?
No. A site audit score measures whether a practice site is structured to be retrieved and quoted. Coverage measures whether assistants actually name the practice. A practice can score 95 out of 100 on the site audit and hold 0 percent coverage on the same day. ClinicAds reports the two numbers separately for that reason.