Most “AI visibility” dashboards measure the wrong thing. A number that goes up after you publish a blog post is not proof — it is correlation wearing a lab coat.

The only honest design: treatment vs. control

For every optimization you ship, we run the same prompt set twice:

  1. Control — the prompt set as it was before your change, answered by the engines today.
  2. Treatment — the same prompt set, after your change is live.

The lift is the difference, on identical questions, with identical wording. Anything else is noise.

The formula

Lift is expressed as the change in citation rate across the prompt set:

lift = citation_rate(treatment) − citation_rate(control)

where citation_rate is the share of prompts in which your brand is named, quoted, or implied, on a fixed, paired prompt set.

A worked example

Suppose a prompt set of 200 questions about your category. Before the change, your brand appeared in 64 of them (control rate = 32%). After a structured claim-verification sprint, it appears in 92 (treatment rate = 46%).

lift = 46% − 32% = +14 percentage points

Across the same 200 questions, on the same day, with nothing else changed — that is a measured effect, not a coincidence.

Citation rate · control vs treatment
Control
32%
Treat.
46%
Δ +14pp Confidence 95%

Four attribution false-positive traps

Even with a clean design, teams fool themselves in four predictable ways:

  1. The novelty spike. A big launch week inflates mentions; you credit the wrong change. Always pair the treatment prompt set with a control set measured the same day.
  2. The seasonal drift. Category interest rises in Q4 for everyone. Lift that “appeared” in December was really the calendar.
  3. The single-engine mirage. One engine changed its defaults; you read it as global lift. Report per-engine, then aggregate.
  4. The repositioning artifact. You reworded the query, not the brand. If the prompt set itself changed, the before/after is not comparable.

Report with confidence, not vibes

A single observation is anecdote. We aggregate across prompt sets and report:

  • The direction of change (cited more / less).
  • The magnitude of change.
  • A confidence interval so stakeholders know how much to trust it.

This is the difference between “we think it worked” and “here is the measured effect, with a confidence level your CFO will accept.”

Why this matters for budgeting

If you cannot attribute lift, you cannot defend spend. Attribution-grade validation turns GEO from a marketing experiment into a measurable line item.