The Wilson confidence interval is a statistical formula tailored to small binary samples (success/failure) that bounds an observed score between a lower and an upper bound. AI Labs Audit surfaces it on every Hero Grade to indicate whether the score is statistically solid or still fragile.
Why not a simple average?
Imagine a brand cited 4 times in 5 prompts (80%) and another cited 80 times in 100 prompts (80%). Both have the same average, but the second is far more reliable. The Wilson confidence interval captures that difference.
How to read it
For each Hero Grade, AI Labs Audit displays for instance:
Score: 78 / 100
Wilson 95% interval: [71; 84]
Reading: with 95% confidence, the true score lies between 71 and 84. The narrower the interval, the more solid the grade. A wide interval signals you should either rerun an audit with more prompts or wait for more historical data.
Why Wilson instead of Wald?
The classic Wald formula (mean ± standard deviation) is numerically unstable when the score is near 0% or 100% and when the sample is small (typical for AEO/GEO audits on 12-24 prompts). Wilson stays accurate in those regions and is the formula recommended by modern statistical standards (Brown, Cai & DasGupta, 2001).
Wilson and AGS
The AGS (Anti-Drift Grading System) computes a Wilson interval on every Hero Grade dimension and on every judged response. This makes fragile scores (interval > 15 points) instantly visible and flagged in the Premium PDF, preserving report defensibility.
Every question asked to ChatGPT without your name in the answer is a competitor recommended instead of you — measured across 6,820 real AI answers.