Full disclosure: unitQ publishes this guide and builds unitQ Support, so weigh the table accordingly and verify against your own evaluation.
A support QA scorecard is a structured rubric for evaluating customer support conversations against explicit criteria, typically accuracy, process compliance, communication quality, and resolution. Agents trust a scorecard when every criterion is observable in the conversation itself, when weights reflect what actually hurts customers, when enough of their work gets reviewed to be representative, and when a disputed score gets a genuine second look. Most scorecards fail on the last two points: they grade a tiny random sample against subjective criteria, then feed the resulting number into performance reviews.
This guide walks through the anatomy of a scorecard that holds up, a seven-step build process, and the coverage question that AI has quietly reopened. For the broader discipline, start with what support QA is.
Why agents stop trusting scorecards
Three failure modes account for nearly all of the resentment QA programs generate.
- 1
Sampling injustice. When reviewers can only grade one or two percent of conversations, a single unlucky ticket can define an agent’s month. The agent knows the sample is not representative. So does the reviewer. The score gets treated as weather, not feedback.
- 2
Subjective criteria. “Was the agent empathetic” is a question about the reviewer’s mood as much as the agent’s work. Two reviewers grading the same ticket should land on the same score almost every time; when they do not, agents learn that scores measure who graded them, not how they performed.
- 3
Consequences without recourse. If a score can affect pay or promotion but there is no published path to challenge it, agents rationally disengage. They stop reading feedback and start managing optics.
Trust is the product of coverage, clarity, and recourse. Design for all three or accept a scorecard that agents comply with and quietly ignore.
The anatomy of a scorecard that earns trust
- 1
Observable criteria. Every line on the rubric should be answerable from the conversation alone. “Did the agent verify identity before discussing account details” is observable. “Did the agent show ownership” is not, until you define the behaviors that count as ownership and list them. A useful test: hand the same conversation to two trained reviewers; if they disagree on a criterion more than occasionally, rewrite it.
- 2
A scale that fits the question. Compliance items should be binary; either identity was verified or it was not. Judgment items (tone, clarity, effort) work better on a short scale, zero to two, than on a ten-point scale that invites noise. Wide scales feel precise and measure nothing.
- 3
Weights that match customer impact. A wrong answer costs the customer more than a missing sign-off phrase, and the weights should say so plainly. When accuracy and script adherence are worth the same points, agents optimize for the points that are easiest to collect, which is exactly the behavior the scorecard was supposed to prevent.
- 4
Auto-fail, used sparingly. Reserve automatic zeroes for genuine red lines: security and privacy violations, abusive conduct, actions that create legal exposure. A scorecard with a dozen auto-fail conditions tells agents that everything is critical, which they correctly read as nothing being critical.
Build one in seven steps
Step 1Start from your support drivers, not a template.
Pull your top contact and escalation reasons and ask what “handled well” means for each. A scorecard copied from a blog post measures someone else’s operation. Understanding what drives your tickets comes first.
Step 2Draft criteria with agents in the room.
Agents know where the rubric will be gamed and where it will be unfair before you ship it. Co-authorship is also the cheapest trust you will ever buy.
Step 3Write pass and fail examples from real tickets.
Every criterion gets at least one anonymized example of meeting it and one of missing it. If you cannot find real examples, the criterion is probably measuring something that does not occur, or something nobody can see.
Step 4Pilot silently for two weeks.
Score real conversations without publishing results or attaching consequences. You are testing the instrument, not the agents. Look for criteria where scores cluster oddly or reviewers diverge.
Step 5Calibrate reviewers on a schedule.
Have every reviewer grade the same set of tickets, then discuss every disagreement until the rubric wording resolves it. Track agreement over time; when it slips, the rubric has drifted or a reviewer has.
Step 6Set the coverage bar honestly.
Decide what percentage of conversations gets evaluated and say the number out loud to the team. If it is two percent, do not present monthly scores as a verdict on anyone’s work. Treat them as a sampled signal with error bars.
Step 7Publish the dispute path.
A named person, a stated turnaround, and a real chance of reversal. Overturned scores are calibration data, not embarrassments; feed them back into step five.
Sampling was a constraint, not a principle
Manual QA samples conversations because human review hours are finite, not because sampling is good methodology. The constraint has changed. AI evaluation can now score every conversation against the same rubric, with human reviewers auditing the AI’s grading rather than grading raw tickets themselves.
This inverts the trust math. Full coverage removes the sampling injustice entirely; no single ticket defines a month when every ticket counts. It also changes what QA finds, because patterns invisible at two percent coverage, a policy that confuses customers on one specific plan, a macro that misfires in one language, show up immediately at one hundred percent.
unitQ Support, unitQ’s AI support QA product, takes this approach: AI scores conversations at full coverage against your rubric, humans audit and calibrate the AI, and the quality signals flow into the same feedback taxonomy as your product data, so a support quality problem that is really a product problem gets routed to the product team (see how to turn support tickets into product insights, coming soon) instead of a coaching doc. If your agents are AI rather than human, the same rubric logic applies with different failure modes; see AI agent QA.
See how unitQ compares on your data
A short demo, run on your own feedback.
Three ways to run a scorecard
Full disclosure: unitQ publishes this guide and builds unitQ Support, so weigh the table accordingly and verify against your own evaluation.
| Spreadsheets + manual sampling | Dedicated QA platforms (MaestroQA, Playvox class) | AI-native QA (unitQ Support) | |
|---|---|---|---|
Typical coverage | A few percent of conversations | Sampled, with workflow tooling to grade more efficiently | Every conversation, AI-scored |
Scoring consistency | Depends entirely on reviewer discipline | Calibration workflows help, humans still vary | Consistent by construction; humans audit the model |
Agent-facing fairness | Weakest; small samples, opaque process | Structured dispute and coaching workflows | Full coverage removes sampling disputes; AI grading needs a visible audit trail |
Product feedback linkage | None | Generally outside scope; QA data stays in QA | Shares a taxonomy with product and review feedback |
Best for | Teams under roughly ten agents | Coaching-centric human QA programs | Scale, consistency, and quality signals beyond coaching |
Capability cells reflect each vendor's published positioning as of August 2026.
When something else is the better choice
Honesty matters more than winning the comparison. If grading tickets by hand is itself your coaching ritual, and your QA leads value that ritual, a manual-first platform in the MaestroQA mold is a defensible choice; the unitQ Support vs MaestroQA comparison (coming soon) goes deeper. If your QA program lives inside a broader workforce management practice covering scheduling and forecasting, a Playvox-class suite bundles those functions together. And if you run fewer than about ten agents, a well-built spreadsheet plus weekly calibration will beat any platform on cost for a long time.
One more limit worth naming: a scorecard measures how conversations were handled, not whether they needed to happen. If the same product defect generates a thousand perfectly handled tickets, your QA score will be excellent and your customers will still be suffering. Pair conversation QA with driver analysis.
FAQ
Ready to see full-coverage QA against your own rubric?
unitQ Support scores conversations at full coverage against your rubric, with humans auditing and calibrating the AI. Take a look and see it against your own rubric.