How to measure CSAT when almost nobody answers the survey
Open your CSAT number and look underneath it. Not the score, the count. In most contact centers a month of ratings is a small fraction of a month of conversations, and the fraction is not random: people rate when they are delighted or furious, and the enormous middle, the customers who got a competent answer and went back to their day, are missing entirely.
That produces a specific set of problems. The score is stable and slightly flattering. It moves for reasons nobody can explain. And at agent level it is close to meaningless, because an agent with a handful of ratings a month has a score that one bad Tuesday can swing by a point.
The three things a low response rate breaks
- Coverage. Whole channels can go unmeasured. Voice is the usual casualty: sending a survey after a phone call means an SMS or an email that arrives when the moment has passed, and the response rate falls further.
- Granularity. Satisfaction per agent, per queue, per issue category and per language are the cuts that would actually change something. Each of them divides an already small sample until the number is noise.
- Silence. The customer who is quietly unhappy and does not rate is the one you most need to hear from, and they are the least likely to reply. Survey based CSAT is structurally deaf to them.
None of that is fixed by asking harder. Sending more survey requests to the same customers reduces the response rate over time and irritates the people who were fine.
What AI scored CSAT does
The alternative is to score the conversation itself. When the room closes, the model reads what was said and assigns a score on the same five point scale, from very unsatisfied to very satisfied, based on sentiment, tone and whether the issue was actually resolved. No rating request goes to the customer.
In our platform it runs across WhatsApp, LINE, Facebook, email, live chat and voice, scores conversations in any language, and lands in the CSAT column of the CX interaction log, filterable by agent, team, channel or date range, next to the sentiment score and the disposition for the same case.
The change that matters is coverage. Every closed conversation has a score, which means the cuts that were noise become usable.
| Question | Survey CSAT | AI scored CSAT |
|---|---|---|
| How satisfied were customers this month | Yes, among those who replied | Yes, across every closed conversation |
| Which issue category makes people unhappiest | Rarely, the sample splits too thin | Yes, by disposition and category |
| How is this agent doing | Only over a long period | Yes, on their own caseload |
| Is our Thai support as good as our English support | Almost never | Yes, scored in the language it happened in |
| What did the customer actually think | Yes. This is the point of a survey | No. This is the limit, below |
What it is not, and this part matters
An AI score is a reading of the conversation. It is not the customer's opinion, and treating the two as interchangeable will eventually embarrass you.
Three limits worth writing on the wall:
- It cannot see what happened afterwards. A conversation that ends with a polite promise reads as satisfied. If the refund never arrives, the customer is not satisfied and the score will never know.
- It reads expression, and expression varies. Customers who are curt by habit, or writing in a second language, or communicating in a culture where complaint is indirect, do not express dissatisfaction the same way. Compare like with like, and be careful reading small differences between markets.
- A contractual number should still come from customers. If your SLA with a client, or your board pack, quotes CSAT, keep the survey for that figure. Replacing a customer reported metric with an inferred one, without saying so, is a reporting problem rather than a measurement upgrade.
Run both, and calibrate
The useful configuration is not one or the other.
- Keep the survey, on a sample. Fewer requests, sent where the answer is worth most, so response rate stops falling.
- Score everything with AI. That is your operational number, the one you cut by agent, category and language.
- Check the two against each other every quarter. Pull the conversations that have both a survey rating and an AI score and see whether they agree. If they diverge in a consistent direction, you have learned something real about either the model's read or the survey's bias, and you can say by how much.
- Use the AI score to find, not to judge. Its best use is a filter: show me every case last week scored one or two. That is a coaching list and a churn list, and the transcript is right there on each case.
One thing not to do, at least not early: put an uncalibrated AI score into a bonus or ranking scheme. The moment a number determines pay, people optimise for the number, and here that means agents writing for the model rather than for the customer. Coach with it for a quarter, calibrate it against real ratings, and decide after that.
Sentiment and CSAT are different columns for a reason
They get conflated constantly. Sentiment is the emotional tone of the interaction, which can be negative throughout a conversation that ends perfectly well: someone reporting a fault is not cheerful, and should not be. CSAT is a judgment about the outcome, closer to whether the customer got what they came for.
Read together they are more useful than either alone. Negative sentiment with a good CSAT is a hard problem handled well, which is often your best agent's caseload. Neutral sentiment with a poor CSAT is the quiet failure that no survey would ever have surfaced.
Common questions
How accurate is AI CSAT compared to a real survey?
The honest answer is that accuracy is a property of your conversations rather than a fixed figure, which is why the calibration step above is in the method. On clear cases, a resolved request or an unresolved complaint, the two agree readily. The disagreements cluster where you would expect: short conversations with little to read, and cases where the outcome happened after the chat ended. Measure it on your own data before quoting a number to anybody.
Does it work on phone calls?
Yes, because a call becomes text. The recording carries a timestamped transcript and a written summary, and the score is read from that in the same way as a chat. This is where the coverage gain is largest, since post call surveys are the hardest to get answered.
Does it work in Thai, Burmese or Bahasa?
Scoring runs on conversations in any language, which is the practical point in this region: a team handling three languages usually has survey data for one of them. Treat cross language comparisons carefully all the same, since expression differs, and read each language's trend against itself rather than against the others.
Do we still need to ask the customer anything?
For a contractual or externally reported number, yes. For running the operation day to day, an AI score on every conversation tells you more than a survey answered by a small and self selected group. Most teams end up sending far fewer requests, and paying much closer attention to the ones they do send.
Where to start
Take last month, filter to the conversations scored one or two, and read ten of them. That is a faster route to a real problem than any dashboard, and it works on day one because the transcript is attached to the case.
The score sits in the same record as everything else: read what case management is for the rest of that record, and setting SLA targets for the other half of the quality picture. Reporting and CX log covers the dashboards over it, and Kai is the AI agent whose own conversations are scored the same way as everybody else's.
To see it against your own conversations, book a thirty minute session.








