Auto QA

QA every conversation. Human and AI.

SentiSum reads 100% of your support conversations — no sampling — and scores each one across every dimension of your QA rubric, with a predicted CSAT and the reason behind every score. Human agents and AI agents, on one scorecard.
Check icon

Independent of your bot vendor

Check icon

Chat and voice, read at source

Check icon

No change to your stack

Trusted by teams at

The problem

Your quality program runs on proxies.

The simple image

/ 01

The 2% sample

Your QA team reads a sliver of tickets and coaches from anecdote. The other 98% of conversations are ungoverned — nobody will ever look at them.

The survey image

/ 02

The survey

CSAT comes from the few who answer — the furious and the delighted. The silent majority, including most bot conversations, never registers at all.

Containment image

/ 03

The containment rate

Your bot vendor's headline metric, measured by your bot vendor. It counts conversations kept away from humans — not problems solved.

Today

With Auto QA

QA coverage

2–5% of tickets, sampled

Check icon

100%, every day

CSAT visibility

The ~5% who answer surveys

Check icon

Predicted for every conversation

The bot's number

Containment, self-reported

Check icon

Resolution, on the human rubric

Coaching

Anecdotes from random tickets

Check icon

Ranked answers, evidence attached

Churn

Found in next quarter's report

Check icon

Flagged mid-conversation

What a first run finds

Containment isn't resolution.

Bot vendors report what stayed in-channel. An audit scores whether the customer got what they came for. Three failure patterns show up in every run — here they are, verbatim.

01 · The needless question

“Where's my driver? The app's said four minutes away for half an hour.”

Bot Please say or key in your nine-digit order number, followed by the pound key.

26% of calls asked for an ID the caller's phone number had already resolved. 66% of them hung up.

02 · The misread that matters

“There's a charge on my card from you that I never made.”

Bot Let's get you back into your account. I've sent a password-reset link to your email.

Quality falls 68 → 15 when intent is misread — and the misreads include fraud, safety anddamage reports.

03 · Churn, announced

“Third damaged box in a row. Pausing was supposed to fix this. Cancel my subscription.”

Bot We'd hate to see you go! Would 40% off your next two boxes change your mind?

81% of customers who stated churn intent were lost anyway — despite being correctlyunderstood.

What you get

The scorecard, the evidence, the fix list.

Live for every conversation, from day one. The first audit is just the first look.

01

The scorecard — see where quality breaks

CQS and predicted CSAT for every conversation, rolled up by intent, team and channel. A Monday-morning glance shows which intent slipped and whose numbers moved.

Quality breaks image

02

The evidence — open the conversation behind any number

Every score opens into the transcript, the dimension-by-dimension verdicts and a plain-language driver. When the scorecard says 41%, you can judge the call yourself in one click.

Understanding

Ownership

Communication

Robotic reply

Resolution

+ pCSAT · sentiment · driver

+ dimensions of your own

Conversation image

03

The fix list — leave with actions, sized by volume

Findings rank by what they'd recover, not by how interesting they are. Most fixes are integrations and routing rules — not a new bot.

Leave actions image
Meet Kyo

Don't dig through dashboards. Ask.

Every score on this page is queryable in plain English. Kyo is the analyst on top of Auto QA — teams ask who needs coaching and on what, why a metric moved, even how a score is calculated. Then they build coaching plans and SOPs straight from the answers.

“Which agents need coaching on understanding, ownership orcommunication — and on what, specifically?”

“Why did CQS coverage drop 32.7% last week when ticket volumebarely moved?”

“Give me the exact rule for how predicted CSAT is calculated.”

“Team health report — this week vs last, quality and pCSAT byteam.”

Real questions from live customer usage, generalised. The last one runs as a standing weekly report.

Meet kyo image
Why Auto QA

Built for the team that answers for quality.

Coverage

QA every conversation — and coach from evidence

Your QA team stops scoring a 2% sample and starts coaching from all of it. Every agent, every conversation, every day — so no complaint, churn threat or compliance risk goes unread, and coaching conversations start from a ranked list instead of an anecdote.

One scorecard

See whether your bot beats your BPO — on the same rubric

Bot analytics score the bot. QA tools score people. Auto QA scores both, identically — so when you decide whether to expand automation, fix the bot, or renegotiate a BPO contract, all three options sit on one scorecard.

Trust

Numbers you can defend — to agents and to the board

Every score carries a plain-language reason and the transcript behind it, computed the same way every time and blind-tested at 88% agreement against an independent model. When an agent disputes a score, you open the conversation. When the board questions CSAT, you show the method.

We publish our agreement rates. Ask your bot vendor for theirs.

Don't just take our word for it. Hear it from our customers.

01
/
04

Frequently asked questions

Answers to common questions about SentiSum's capabilities, set up, and how we help reduce churn

Frequently asked questions

Answers to common questions about Sentisum’s capabilities, set up, and how we help reduce churn

Quality Agent

How is this different from our bot vendor's analytics?
Your vendor measures containment, and containment is the metric this page exists to correct. Auto QA scores resolution, independently, on the same rubric as your human agents — and predicts CSAT for the conversations no survey reaches. Structurally, a vendor cannot mark its own homework. We have no bot to defend.
We already test our agent before every release.
Keep doing that. Pre-release simulation and regression testing are engineering QA, and they answer a different question: does the agent behave as designed? Auto QA answers the one that follows: does the design resolve real customers? Every failure pattern on this page came from an agent that passed its tests — the needless question, the misread incident, the unactioned cancellation aren't bugs. They're design gaps, and they only show up in production conversations, which is where we read.
What does it measure, exactly?
The core rubric scores understanding & relevance, ownership & proactivity, communication quality, robotic-reply detection and issue resolution, rolled into a composite Conversation Quality Score — and it’s config-driven, so teams add dimensions of their own on top. Alongside it: a predicted 1–5 CSAT and a start-to-end sentiment journey. Reported per intent, per channel, per team, and every score carries a plain-language driver reason plus the conversation IDs behind it.
How do I know the scores are right?
We re-score a stratified blind sample with a stronger, independent reference model and publish the agreement: 88% on support-request classification, 84% on predicted CSAT within one point. Lower-confidence labels are treated as directional and reviewed by an analyst before reporting. What counts as “resolved,” what's included, and what's excluded are written rules you can inspect and change.
Which bot platforms does it work with?
Any chat or voice agent, read at source. Recent audits have scored chat and voice bots from different vendors in a single report — if your customers talk to it, it can be scored. There is no integration project on your side.
Does this replace my QA team?
It replaces their sampling, not their judgment. Today they read 2–5% of conversations chosen mostly at random; that time moves to coaching, calibration and the edge cases the scoring flags for human review. Same headcount, materially more coverage.
What happens to our customer data?
Conversations are read at source over read-only access, used solely to produce your scores and findings, and never used to train models for other customers. Retention and redaction follow your policies, and security documentation and a DPA are available on request.
What does a first audit involve?
You grant read access to a month of conversations; we return a findings report in the format shown on this page — the resolution breakdown, the failure patterns with transcripts, churn-risk conversations, and a ranked fix list. If it tells you nothing your dashboard didn't, walk away.