Back to Blog

Contact Center QA Metrics That Matter: A Practical Guide for COOs and QA Leads

6 min readBy The Caller.ee Team
Quality AssuranceCoaching
#QA Metrics#Scorecards#Agent Performance#Analytics

When QA runs on manual sampling, metrics are mostly decoration. If you're scoring 20 calls a month out of 5,000, your "average quality score" has a margin of error wide enough to drive a truck through, and every number downstream of it inherits the problem.

Score every call — or even most calls — and the same metrics suddenly mean something. Now the question becomes: which numbers should a COO or QA lead actually look at, in what order, and what should each one trigger?

This is a practical tour of the call center QA metrics Caller.ee computes from your scored calls, how to read each one, and how to string them into a weekly ritual that turns numbers into coaching. We'll close with the metric's honest limits — what a QA score can't tell you.

Average Quality Score: Your Baseline, Not Your Goal

Every call analyzed in Caller.ee gets an overall score from 0 to 100%, produced by your quality assurance scorecard — the Analysis Template whose sections and weights you defined. The average of those scores is the single number most executives will ask for.

Use it for what it's good at: establishing a baseline and detecting movement. If your campaign averaged 74% for three months and drops to 66% this week, something changed — a new hire wave, a script revision, a product issue generating harder calls.

Don't use it as a target in isolation. An average is a blender: it happily mixes ten excellent calls with ten terrible ones and reports "fine." The metrics below exist to un-blend it.

Fail Rate: The Number That Should Alarm You

Caller.ee tracks the percentage of calls scoring below 70 — the fail rate. This is usually more actionable than the average, because customers don't experience your average. Each customer experiences one call, and a failed call is a concrete bad experience, a compliance risk, or a lost sale.

Two operations can share a 78% average while one has a 5% fail rate and the other 20%. The second one has a fire. Watch fail rate weekly, per campaign; a rising fail rate with a stable average means your worst calls are getting worse, which is exactly the pattern averages hide.

Score Distribution: The Shape of Your Team

The score distribution shows how calls spread across score bands. Three shapes to recognize:

  • Tight cluster, high center. A trained, consistent team. Your process works; protect it.
  • Two humps. Two populations — commonly veterans and a recent hire class, or two sites executing differently. Coach the populations, not the average.
  • Long left tail. Mostly good with a stubborn stream of very bad calls. Go read that tail — Caller.ee's lowest-scoring calls list takes you straight to them, and the executive summaries tell you in seconds whether it's an agent, a call type, or a broken process.

Weakest Criteria: The #1 Coaching Signal

If you look at only one screen a week, make it this one. Because every scorecard section is scored separately, Caller.ee can rank your criteria worst-first across all analyzed calls: which specific behaviors fail most often.

This is where QA stops being surveillance and becomes diagnosis. "Quality is down" is a mood; "'proposed a concrete next step' is our lowest-scoring criterion, failing across most of the team" is a training plan. And when a criterion fails broadly — not just for a few agents — the root cause is almost never the agents. It's the script, the training, or an unrealistic expectation. Fix it once, at the source, instead of coaching the same gap into twenty people one by one.

Agents to Coach: Where the Next Hour Goes

Coaching time is the scarcest resource a team lead has. The agents-to-coach view ranks agents by lowest average score, so one-on-one time flows to where it changes the number most.

Read it alongside per-agent section scores: two agents with identical 68% averages might need opposite conversations — one is weak on discovery questions, the other on closing. Since results also carry the campaign context, you can check whether an agent is struggling everywhere or only on one campaign's material — a distinction that separates a skills gap from a knowledge gap.

One caution: rankings are for allocating help, not for public shaming. Post the leaderboard on the wall and agents will optimize for the scorecard's letter rather than the customer. Use it privately, pair it with the actual calls, and coach from evidence.

Campaign Performance: Comparing Like With Like

Because calls in Caller.ee live inside a Clients → Services → Campaigns hierarchy, you can compare quality within fair boundaries. A cold outbound campaign and an inbound support line should never share a benchmark — different scorecards, different difficulty, different goals.

Campaign comparison answers the operator questions: Which client engagement is healthy and which needs intervention? Did the new campaign ramp to target quality by week three or is it stuck? Where do we point our best team leads next month?

Score Over Time: Did Anything We Did Work?

Every intervention — new script, training day, coaching push — is a hypothesis. Score over time is where hypotheses go to be tested. Ran objection-handling training in week 2? The objection-handling criterion's trend in weeks 3–6 tells you whether it stuck. No movement means the training didn't transfer, and you've learned that in a month instead of a quarter.

Trends also catch slow decay: quality tends to erode gradually — script drift, waning attention after onboarding — and a gentle six-week slide is invisible in any single week's numbers but obvious on a trend line.

Handle Time and Volume: Context, Not Verdicts

Caller.ee also tracks average handle time and call volume. Treat these as context for quality numbers, not as goals. Short calls aren't automatically efficient — they may be agents rushing past discovery (your weakest-criteria view will say so). Long calls aren't automatically thorough. And a quality dip during a volume spike reads very differently from one during a quiet week. When you want to go deeper, everything exports to CSV.

The Weekly QA Ritual

Here's how these metrics compose into a repeatable 45-minute weekly session for a QA lead or ops manager:

  1. Scan the top line (5 min). Average score, fail rate, score over time — per campaign. Note anything that moved.
  2. Open weakest criteria (10 min). Identify this week's worst one or two criteria. Decide: systemic (fix script/training) or localized (coach individuals)?
  3. Read the tail (10 min). Open the lowest-scoring calls. Read summaries; open the transcript of the two or three worst. This keeps you honest about what the numbers mean in human terms.
  4. Pick coaching targets (10 min). From agents-to-coach, choose two or three agents. For each, pull one weak call and one strong call as session material.
  5. Log one process action (10 min). Every week, one systemic fix: a script line, an FAQ update to the knowledge base your templates are grounded on, a training note. Mark it, so next month's trend line can judge it.

Small, boring, weekly — that cadence beats a heroic quarterly QA audit every time.

What a Score Can't Tell You

A QA score measures conformance to your standard — did the agent do the things your scorecard says a good call contains? It does not measure customer satisfaction. A customer can leave delighted from a call that broke your script, and furious after one that followed it perfectly. Caller.ee doesn't run CSAT surveys or track resolution outcomes; if you measure satisfaction elsewhere, treat the two as complementary lenses, and be suspicious when they diverge — a rising QA score with falling satisfaction usually means your scorecard is testing the wrong things. The scorecard is a hypothesis about what a good call is. Revisit it quarterly.

Put Your Own Numbers on the Board

The fastest way to make this concrete is with your own calls. Caller.ee's free tier — live now in public beta — gives you 45 transcription minutes, 30 calls, and 15 analysis runs a month, enough to score a real sample from one campaign and see your baseline, fail rate, and weakest criteria by the end of the day. Start at https://app.caller.ee; if you need more room, email info@sumgrey.com and we'll raise your limits.

See your own calls scored

Caller.ee is free during our public beta — transcribe, analyze and score real calls on European infrastructure, no credit card required.

Start Free