AI visibility metrics that matter

Four metrics cover AI visibility properly: Share of AI Answer (the headline — the percentage of tracked questions where you are cited), per-engine citation rate (where you are weak), competitor win rate (who beats you on the questions you lose), and answer-source mix (which domains engines cite in your category — effectively your to-do list). Everything else is commentary on those four.

Last updated 2026-07-28

Which four AI visibility metrics are worth tracking?

These four, in this order. The first tells you how you are doing, the second where, the third against whom, and the fourth what to do about it.

  • Share of AI Answer — cited results ÷ completed results across your prompt set. Read the 30-day trend, not today's number. Full definition and formula in what is Share of AI Answer.
  • Per-engine citation rate — the same score split by engine. A 50% ChatGPT / 10% Gemini split is an action plan, not a curiosity: it says your problem is Google-index visibility rather than AI presence generally.
  • Competitor win rate — on the prompts where you are not cited, who is? The domains that repeatedly win your lost prompts are your real competitors in AI answers, and they are frequently not the competitors on your slide deck.
  • Answer-source mix — the third-party surfaces (directories, communities, publications, video) engines cite for your category. Presence there is usually cheaper than outranking anyone, and it is the input to a realistic AEO plan.

How do I read the trend without fooling myself?

Treat every daily run as a sample, not a score. Grounded answers are stochastic: the same prompt can cite you today and not tomorrow with nothing having changed on the web, so a single-day movement carries almost no information.

Three rules make the trend honest. Read weekly averages rather than daily points. Compare like with like — a full sweep over 120 prompts is not comparable to a core run over 15, so keep the two series separate. And hold the prompt set constant; if you must add prompts, note the date and expect the level to shift, because you changed the denominator. AI Visibility Tracker reports the change against the previous completed run in percentage points for exactly this reason: a delta is honest about what it is, where a raw number invites over-reading.

Which AI visibility metrics should I ignore?

Mostly anything without a fixed denominator or with a sample too small to support the claim being made.

  • "AI mentions" with no denominator — 12 mentions of what? Without a fixed prompt set there is no rate, no trend, and no way to compare two weeks.
  • One-off screenshots — answers vary run to run, so a single flattering sample proves nothing, and neither does a single bad one.
  • Sentiment scores on tiny samples — with a few dozen answers a week, classification noise swamps any real signal. Sentiment becomes meaningful at volumes most brands are not running.
  • Estimated "AI traffic" or "AI referrals" projected from visibility scores — the conversion from citation to visit is unmeasured and varies by engine and query type. Report what you measured, not a model of what it might be worth.
  • Composite "AI visibility scores" that blend several metrics into one index — they are unfalsifiable, they move for reasons you cannot decompose, and they cannot be audited by anyone outside the vendor.

How many prompts do I need for the numbers to mean anything?

Ten to fifteen core prompts is enough to start; the useful precision arrives from repetition over time rather than from a huge prompt set on day one. Across five engines, fifteen prompts already produce 75 results per run and roughly 500 a week, which is enough to distinguish a persistent shift from ordinary variance.

Beyond that, add prompts to cover intents rather than to inflate the count: discovery ("best X for Y"), comparison ("A vs B"), alternatives ("A alternatives"), and brand checks ("is A any good"). A prompt set that covers four intents at fifteen prompts is more informative than one hundred paraphrases of the same discovery question — and materially cheaper to run daily.

Where should this data live?

Somewhere you control, in a form you can query. A visibility prompt set is a compact description of your commercial strategy — the questions you believe your buyers ask and the competitors you consider real — and the result history is the only asset in this discipline that cannot be recreated after the fact.

AI Visibility Tracker keeps everything in a local SQLite database on your own machine, including the full raw response from every engine call, so any historical number can be audited back to the answer that produced it. API keys are encrypted with the operating system keychain and never written to that database in plaintext; there is no telemetry and no vendor server in the path. If that trade-off matters to you, how AI Visibility Tracker compares is honest about what you give up for it.

Frequently asked questions

What is a good number of prompts to track?

10–15 core prompts run daily, plus a longer extended set swept weekly. Cover four intents — discovery, comparison, alternatives and brand checks — rather than writing many paraphrases of the same question. Fifteen prompts across five engines already yields around 500 results a week.

Should I track competitors as well as my own brand?

Yes — competitor data is the actionable half. On every prompt where you are not cited, recording which domains were tells you where engines already look for your category, which is a ranked to-do list rather than a scoreboard.

Can I export AI visibility data for my own reporting?

In AI Visibility Tracker, yes: results export to CSV and PDF, and the app can expose its own data to your other AI tools over a local MCP server bound to 127.0.0.1. Because storage is a plain SQLite file on your machine, the data is queryable directly as well.

Related guides