The inter-judge GRC (Generalized Reliability Coefficient) is a consistency metric across the 5 LLM judges of AGS. The closer to 1, the more the judges agree; the closer to 0, the more the final score is debated, and therefore fragile.
Why measure judge agreement?
AGS mobilises 5 LLM judges (Claude, GPT-4, Gemini, Llama, Mistral). When all 5 agree, the final score is solid. When 3 say A and 2 say F, the mean is C — but that mean hides a disagreement that should be surfaced. That is the role of GRC.
How to read it
- GRC = 1.0: perfect agreement. All 5 judges give the same grade.
- GRC ≥ 0.8: strong agreement. The grade can be defended as-is.
- GRC between 0.5 and 0.8: moderate agreement. The grade is still usable but the Wilson interval will be wider.
- GRC < 0.5: strong disagreement. AI Labs Audit flags the dimension as “contested” in the Premium PDF and recommends rerunning an audit or prioritising that dimension in the action plan.
Tie-in with Hero Grades
GRC is computed for every Hero Grade dimension (citations, position, sentiment, coverage, authority, perception). It is displayed alongside the Wilson interval in the dashboard and PDF, allowing GEO consultants to defend each point rigorously in front of a client.
Every question asked to ChatGPT without your client's name in the answer is a competitor recommended instead of them — measured across 6,820 real AI answers.