Skip to main content

Every company is deploying AI support... few are grading it

Sep 23, 2026
By Christian Wiklund
Co-founder & CEO at unitQ
Every company is deploying AI support… few are grading it

If you’ve deployed an AI support agent, it’s likely answering your customers right now. They’re resetting passwords, explaining charges, processing refunds, and telling people why their order is running late. The quality program built to catch a bad support interaction was designed for humans. And it still reviews a small sample of what goes out the door.

That worked well enough when a team checked another team’s work. Spot-check a few conversations a week and you had a rough read on the whole. An AI agent breaks that math. It answers at a volume no weekly sample can keep pace with, so the spot-checks that once covered a meaningful share now cover a rounding error.

Most teams are still auditing their bots the way they audited people, which no longer tells them much. And customers are unforgiving of the misses. In a HubSpot and SurveyMonkey study of 15,000 consumers, just over half said they actively dislike or hate AI in service interactions, and 82% would rather deal with a human even when wait times are identical. 

What AI reports and what customers feel are two different stories. Only one of them shows up on a dashboard.

Support QA was broken before AI arrived

Traditional quality assurance runs on sampling. A reviewer pulls a fraction of tickets each week, often only a few percent, and scores them against a rubric assuming the rest look similar. That foundation is already shaky. The failures that hurt a support operation most—a refund issued for the wrong amount or a policy exception granted to the wrong customer—don’t happen often, which is exactly why a small sample misses them. A sample that small is statistically unlikely to surface a problem that shows up in one ticket out of hundreds. And when it does catch one, there’s no way to know how many more slipped past unread.

One wrong answer, ten thousand times

A human agent who gives a wrong answer gives it to one person. An AI agent gives it to everyone who asks the same question until someone notices and fixes the underlying content. Systemic errors like this can run for weeks before anyone notices, leaving one bad response to become thousands of bad responses at machine speed.

AI quality also drifts silently. A knowledge base update, a product change, a reworded policy—any of them can degrade how the AI responds, with no alert and no obvious signal. Teams usually find out the way Cursor and Klarna did.

In April 2025, a developer asked Cursor’s support why he kept getting logged out. The company’s AI support bot, signing its emails “Sam,” explained that Cursor only works on one device per subscription. That policy didn't exist. The bot invented it. Users took it at face value and cancelled before anyone at the company noticed, and a co-founder had to post on Reddit to say there was no such rule. The bot didn’t even hallucinate consistently. Some users were told the fake policy, others weren’t, so customers comparing notes couldn’t tell what was real.

The quiet version is worse, because you don’t get a Reddit thread to warn you. Klarna’s AI assistant handled 2.3 million conversations in a month and cut average handle time from 11 minutes to under two. By its own metrics, a win. Then the CEO told Bloomberg the company had “focused too much on cost,” that “the result was lower quality,” and began rehiring. The metrics that declared victory measured speed and average satisfaction, but said nothing about whether hard problems were getting solved. The quality failure stayed invisible until it wasn’t.

Their story is already playing out across the industry. Gartner projects that agentic AI will autonomously resolve 80% of common customer service requests by 2029, and also predicts that 50% of organizations who cut support headcount citing AI will rehire by 2027. The deployment wave and the walk-back wave are happening at the same time.

So who’s actually grading AI?

In a lot of deployments, the answer is uncomfortable: the AI vendor grades its own AI. 

Yup, the same system that generated the response also reports on how good the response was. That’s the watchmen watching themselves, and it’s a shaky foundation for a function whose entire job is catching failure.

The tooling gap runs deeper. Most support teams now run a hybrid model where AI handles tier-one volume, with humans taking escalations and nuanced cases. A QA tool that only grades humans or only grades its own bot leaves half the operation unmeasured. You end up managing two quality standards that can’t be compared, for one customer who doesn’t care which side answered.

What real AI support QA looks like

If you’re going to trust AI to talk to customers, the QA has to be built for that reality. 

Score every conversation, not a sample

A model can read and grade a conversation in seconds, which makes grading every conversation attainable for the first time. Full coverage turns QA from a rear-view sample into an early-warning system that catches expensive failures before they repeat.

Grade AI and humans on the same rubric

One customer experience deserves one quality standard. Scoring bot and human conversations against the same scorecard is what makes the numbers comparable and the handoffs visible.

Judge whether the issue got resolved

Grading needs a shared definition of “good.” Here's the rubric every AI support conversation should be scored against:

5 things every AI support conversation should be graded on 
1. Accuracy against policy: Is the answer factually correct and consistent with your actual policies?
2. Resolution: Did the customer’s problem get solved, or did the conversation just end?
3. Escalation and handoff: When the AI hit its limit, did it route the customer to a human cleanly?
4. Tone and empathy: Did the response match the moment?
5. Compliance: Did the AI follow the required steps and disclosures?

Of the five, resolution is the one most programs quietly drop, because it’s the hardest to measure and the easiest to fake. AI is fluent even when it’s wrong, so a conversation can look great on tone and still leave the customer's problem sitting there. Grade against outcome and policy, not politeness.

Close the loop

A score nobody acts on is wasted. SQM Group found that 83% of customer service agents believe their QA program doesn’t actually help them improve customer satisfaction, usually because scores land in a monthly report that never reaches coaching. The point of grading every conversation is to route what’s broken to whoever can fix it, whether that’s the content owner, the model, or the human team. Then confirm it actually improved.

The most honest grade comes from outside the system

Handle time, deflection rate, and self-reported CSAT all measure the AI on the terms of the team that deployed it. They’re the metrics that made Klarna look successful right up until customers revolted. The grade that actually protects the experience comes from a different place: What customers themselves say and experience, measured independently of the system being graded, across every channel where they tell you.

That independent read is what most AI support programs are missing. Not a small sample, but a full-coverage view of quality grounded in the customer’s real experience. It’s the problem unitQ Support QA was built for: grading AI and human support against what customers actually report, independent of the system doing the answering.

The question worth asking

Everyone in support is asking whether AI can handle the work. Almost nobody is asking whether they’d notice if it stopped handling it well. The AI is already answering your customers. The only real question is whether it’s answering them well enough, and whether you have any way to know. Grade AI like it’s talking to your customers. Because it is.

Frequently asked questions

Newsletter

Quality insights, in your inbox

Join product and CX leaders getting the latest on customer intelligence, once a month.