Calibration Sessions: Getting Your QA Scorers to Agree
If two reviewers score the same chat differently, your scores are noise. Calibration fixes that.
Here is a quick test of your quality program. Take one conversation, hand it to two of your reviewers separately, and compare their scores. If the numbers don't match, you don't have a measurement system. You have two opinions wearing a rubric.
QA calibration is how you fix that — the practice of aligning your scorers until the same conversation gets the same score no matter who reads it.
Why QA calibration matters more than the rubric
A great rubric applied inconsistently is worse than a mediocre one applied uniformly. If a rep's score depends on which reviewer pulled their conversation, the score measures the reviewer, not the rep. Coaching built on that is coaching built on noise.
The technical name for what you are chasing is inter-rater reliability: the degree to which independent scorers agree. When it is high, a score means something — you can trust trends, compare reps fairly, and coach with confidence. When it is low, every conversation about performance turns into a debate about the grader, and reps learn to discount the whole program.
Calibration is the routine that drives that agreement up and keeps it there. It is not a one-time setup. Scorers drift, edge cases pile up, and language gets reinterpreted. A program without regular calibration quietly decalibrates.
How to run a calibration session
The mechanics are simple, and the discipline is in following them honestly.
- Pick shared samples. Choose 5–8 real conversations that span easy, hard, and genuinely ambiguous. The ambiguous ones do the real work — easy conversations everyone scores the same anyway.
- Score blind, independently. Each reviewer scores all of them alone, without seeing anyone else's marks. Blind is the whole point; if scorers can peek, they anchor to each other and you learn nothing.
- Surface the deltas. Put the scores side by side and find where they diverge. Ignore the categories everyone agreed on — go straight to the disagreements.
- Discuss each delta to root cause. For every gap, ask why. Usually it is one of two things: a reviewer misread the transcript, or the rubric criterion is vague enough to support both readings.
- Fix the cause, not the score. If it was a misread, that is a coaching moment for the reviewer. If it was vague language, rewrite the criterion on the spot so the next ambiguous case resolves the same way for everyone.
Every disagreement is a gift. It points at the exact criterion that's too fuzzy to be fair — and now you get to fix it.
Read the deltas like a diagnosis
Not all disagreements mean the same thing. Learning to tell them apart is what makes calibration efficient.
| What you see | Likely cause | Fix |
|---|---|---|
| One reviewer way off, others tight | Reviewer misread or has a personal standard | Coach the reviewer |
| Two clean camps, split evenly | The criterion supports both readings | Rewrite the criterion |
| Everyone scattered | The category itself is poorly defined | Rebuild the category |
| Tight agreement everywhere | You're calibrated on this slice | Move to harder samples |
The pattern tells you whether you have a people problem or a rubric problem, and those have different fixes. Spending a session coaching reviewers when the real issue was vague language just moves the disagreement to the next conversation.
Keep it a habit, not an event
Calibration decays. A team that calibrates once at launch will drift within a quarter as new edge cases and new hires accumulate. The teams with reliable scores run short, regular sessions — every couple of weeks, 30 minutes, a handful of fresh samples — rather than one heroic offsite a year.
Track the agreement rate over time. If it is climbing or holding, your program is healthy. If it is slipping, your rubric is aging or your roster has changed, and it is time to recalibrate before the scores lose meaning.
What this looks like with AI in the loop
When an AI agent scores every conversation against your rubric, calibration changes shape but does not go away. Now you are calibrating your reviewers against each other and against the AI's proposed scores — and the deltas are just as diagnostic. Where reviewers consistently overturn the AI on a category, the rubric guidance for that category needs sharpening, for the humans and the AI alike. The receipt behind each AI score makes the disagreement easy to inspect: you can see exactly which transcript lines it scored on.
Closing
Calibration is the unglamorous routine that makes every other part of your quality program trustworthy. Blind scoring, honest discussion of the deltas, and a rubric that gets sharper each session — that is the whole craft, and it pays off every time you act on a score.
BearScope scores every conversation on your rubric with the transcript evidence behind each mark, so reviewers calibrate against real evidence, not memory — and corrections feed straight back into the standard. To see calibrated, full-coverage scoring on your conversations, book a walkthrough, read about how the product works, or learn how we keep it auditable.
See it on your own conversations.
Bring your busiest day. We'll score every conversation in it.
Book a walkthrough →