Quality & Coaching

The QA Scorecard: A Field Guide for Support Leaders

What belongs on a support scorecard, and what's just noise. A practical reference.

Most support teams have a QA scorecard. Most of those scorecards measure the wrong things, score them on the wrong scale, and end in a number nobody acts on. A good scorecard is not a longer checklist — it's a tighter one, built so every mark points at something a person can change on their next conversation.

This is a field guide to what belongs on a QA scorecard and what is just noise. Use it to audit the one you have.

What a QA scorecard is actually for

A QA scorecard exists to answer two questions: did this conversation resolve the customer's problem, and would you be comfortable if every conversation looked like this one? Everything on the card should serve one of those. If a line item doesn't change the answer to either question, it's decoration.

The most common failure is the everything-bucket scorecard — twenty checkboxes covering greeting, sign-off, grammar, formatting, and a dozen micro-behaviors that have nothing to do with whether the customer left helped. You can score 95% on that card and still lose the customer. Cut hard. A scorecard that measures resolution and risk beats one that measures etiquette.

If a line item can't change the outcome of the conversation, it doesn't belong on the scorecard.

The five parts of a working scorecard

A scorecard that drives behavior has five parts, in this order.

  1. Categories — the four to six things you actually care about. For most support teams: accuracy, resolution, tone, setting expectations, and (if relevant) compliance. Name them as behaviors, not vibes.
  2. Weights — not every category matters equally. A factual error that misleads a customer should sink the score in a way a missing sign-off never does. Weight accordingly.
  3. Evidence links — every mark points at the exact transcript line that earned it. A score with no evidence is an opinion, and reps argue with opinions.
  4. A verdict — one plain-English sentence that says what happened: resolved correctly, warm tone, but promised a refund timeline the policy doesn't support. The verdict is what the rep remembers.
  5. A coaching note — the single behavior to change next time, written so the rep can repeat it back.

Drop any of these and the card weakens. Drop evidence links and reps stop trusting it. Drop the coaching note and nothing changes.

Categories and weights, side by side

Here's a starting scorecard you can adapt. The point isn't these exact numbers — it's that the weights reflect what actually hurts a customer.

CategoryWhat it measuresWeightAuto-fail?
AccuracyWas the information correct and complete?30%Yes, if it misleads
ResolutionDid the customer's problem actually get solved?30%No
Setting expectationsWere timelines and next steps honest?15%No
ToneDid the reply acknowledge and respect the person?15%No
ComplianceWere required disclosures and policy followed?10%Yes, if breached

Notice two things. Accuracy and resolution together are 60% — they're the job. And two categories carry an auto-fail: a confidently wrong answer or a compliance breach should cap the score no matter how warm the tone was. A weighted model with critical-fail gates beats a flat average, because a flat average lets a great greeting paper over a dangerous mistake.

What to leave off

Resist the urge to score:

  • Grammar and formatting, unless it actually obscured the answer. Polish the templates instead.
  • Greeting and sign-off wording, unless your brand depends on it. It rarely changes resolution.
  • Anything you can't link to a transcript line. If you can't point at it, you can't coach it.
  • Handle time as a quality measure. It's an operations metric. Mixing it into quality pressures reps to close fast, not close right.

A leaner card gets scored more consistently, which makes the data trustworthy enough to coach from.

Make the score lead somewhere

A scorecard is the start of a coaching loop, not the end of a review. The verdict and coaching note are what turn a number into better work next week. That's why the card should produce one specific behavior to fix — not a 78 and a shrug. For more on closing that loop, see coaching from QA scores, and for the choice between flat and weighted scales, see pass/fail vs. weighted scores.

In BearScope, every conversation — from your reps and from your AI agents like Jenny — is scored on your own rubric, with each mark linked to the transcript line that earned it and a plain-English verdict on top. The receipt behind every score means a rep can see exactly why, and a leader can audit it. See how scoring and review work, or book a walkthrough to build your scorecard live.

The best scorecard is short, weighted toward what hurts customers, and ends in one thing to change. Audit yours against that.

See it on your own conversations.

Bring your busiest day. We'll score every conversation in it.

Book a walkthrough

Keep reading