Last updated
Grace Cowan
Customer service

Anyone who has sat through three AI customer service demos in a week will recognise the feeling. Every vendor arrives with a number. One reports 70% automation, the next 60% resolution, the third 45% deflection. All three sound like the same claim, all three sound good, and by the third meeting it is genuinely difficult to remember which was which.
None of those vendors is lying. The numbers are real. The trouble is that they are measuring four different things, and the industry uses the words as though they were interchangeable.
That is a more expensive problem than it looks. The gap between "the AI handled it" and "the customer's problem was solved" is where service budgets quietly disappear, and it rarely shows up in the reporting. So it is worth spending twenty minutes on the definitions before you spend anything else.
This article covers what each of the four numbers actually measures, how each one gets inflated, how the major vendors define resolution when they bill you for it, and how we define it ourselves.
Why two vendors can report the same performance and very different numbers
Every percentage has a denominator, and the denominator is where the argument lives.
If a vendor's AI only sees conversations that a routing rule sends it, and that rule sends it the easy ones, its resolution rate will look extraordinary. If another vendor's AI sees every inbound conversation including the furious ones, the same underlying quality produces a much worse number.
Then there is the question of what counts as a success. Some vendors count a conversation as resolved if no human joined it. Others count it if the customer did not come back within a set window. Neither of those is the same as the customer getting what they needed.
So when you are handed a percentage, two questions do most of the work. What was the denominator, and what counted as a win?
The four numbers, and what each actually tells you
Here they are side by side. The final column is the one worth dwelling on.
What it measures | What it does not tell you | How it gets inflated | |
|---|---|---|---|
Involvement | The share of conversations the AI touched at all | Anything about quality | Route only simple queries to the AI |
Containment | The share where no human joined | Whether the customer was helped | Make reaching a human difficult |
Resolution | The share where the problem was actually solved | Whether the customer will return about it | Define resolution as "no human joined" |
Automation rate | The share of work handled without a person | What kind of work it was | Count acknowledgements and menu selections as handled |
Read across that table and one thing stands out. Three of the four can be improved without helping a single customer more than you do today. Only resolution resists that, and only if it is defined honestly.
Why containment is the number most likely to mislead you
Containment is the figure most often put in front of buyers, and it is a routing outcome dressed as a service outcome.
A contained conversation is one where nobody from your team joined. That describes a customer who got a good answer and left satisfied. It also describes a customer who asked three times, got nowhere, could not find a way to reach a person, and gave up.
Both of those are contained. One of them is a resolution and the other is a customer who is now telling colleagues that your support is useless.
This matters more than it sounds, because containment responds to exactly the wrong lever. Make your escape hatch harder to find and containment goes up. Bury the "talk to someone" option and containment goes up again. The metric improves as the experience gets worse, which is the definition of a bad measure.
Gartner found in 2024 that only 14% of service issues are fully resolved in self-service, rising to 36% for issues customers themselves rated as very simple. Set those figures against a vendor's 70% containment claim and the distance between the two is the space where customers are quietly giving up.
Silence is not consent. If your reporting cannot tell the difference between a solved problem and an abandoned one, it is not telling you about service quality at all.
How the major vendors define resolution
This has stopped being an academic question. Resolution definitions are increasingly what you are billed against, not just what appears on a dashboard, so the wording in a contract matters more than the number in a deck.
Zendesk bills on what it calls verified resolutions. The interaction has to be handled by AI without human involvement, and a separate language-model evaluation has to confirm the customer's issue was satisfactorily resolved. Interactions where AI contributed but a human finished are classed as assisted escalations. Interactions the AI handled but the evaluation could not confirm are classed as contained resolutions. Neither consumes your resolution allowance.
That is a more careful model than most, and it is worth saying so. It also means an unpublished language model is deciding what you pay for, with no published accuracy figure for the evaluator and no documented appeal process. You can compare the details on their pricing page, and we cover how the model works alongside other options in our comparison of Zendesk alternatives.
Intercom charges $0.99 per resolution through Fin, with an outcome definition broader than classic support resolution. It includes procedure handoff and qualification, which are useful things for an AI to do but are not the same as answering a customer's question.
Neither approach is wrong. Both are incomparable with each other, and with anyone else's.
How Cue defines these terms
It would be a bit rich to spend an article asking you to interrogate other people's definitions without publishing our own. So here they are.
Involvement. Our AI is involved from the start. It is the first interaction rather than a fallback after a human has triaged the queue. That means involvement is not a number we can be impressive about, because the answer is effectively always. It also means our resolution figures are calculated across everything that comes in, not a filtered subset.
Containment. Passed to a human agent, or not passed to a human agent. A binary, and a routing fact.
Resolution. The conversation was resolved. This is deliberately not defined as "no human joined". It means what the AI was actually able to resolve, which is a smaller number and a more useful one.
That distinction is the whole reason this article exists. We define resolution more narrowly than the vendors billing on it, and we would rather report a lower figure that means something.
An example of what the narrower definition produces in practice: Payflex resolved 82% of customer queries with Cue AI Agents without a human joining the conversation. Note the full wording, because it is doing real work. It is a containment-and-resolution figure for one customer with one query mix. It is not a Cue-wide resolution rate, and it is not a benchmark for your business.
Escalation is not one number
Most reporting treats escalation as a single failure bucket. That collapses three quite different situations.
Our AI Agents return a reason when they hand a conversation over, and the reason changes the interpretation:
Reason | What happened | Is it a failure? |
|---|---|---|
| The customer asked for a person | No. A customer entitled to ask for a human asked for a human. |
| The AI recognised it could not resolve this and that a person was needed | No. This is the system working correctly. Knowing your limit is the desired behaviour. |
| Something went wrong | Yes. This one is a genuine failure. |
Only one of the three is a defect. If your reporting cannot separate them, a rising escalation rate tells you nothing about whether anything is wrong. It may even be telling you that your AI has got better at recognising what it should not attempt.
This is also why the handover itself matters more than the rate. A conversation that reaches a person with the history, the context and the reason attached costs a fraction of one that starts again from nothing. We look at how connected service reduces repeat contact separately, because that is where most of the recoverable cost sits.
The measure the category has not solved
Every number above describes conversation handling. None of them describes whether the conversation protected the relationship.
Did the customer stay? Did they buy again? Did the resolution prevent a cancellation, or simply close a ticket?
This is the open problem in service measurement, and it is not close to settled. Connecting a resolved conversation to retained revenue requires an attribution model that survives scrutiny, and the industry does not have an agreed one.
So it is worth asking, and worth listening carefully to the answer. Ask how a vendor connects a resolved conversation to a retained customer, then ask to see the attribution model rather than the dashboard. A confident number here deserves more scepticism than anywhere else in the stack.
Questions to ask any vendor about their numbers
Six questions, and they take about five minutes.
What was the denominator? Every conversation, or a routed subset?
What counted as resolved, in one sentence?
Who or what decided it was resolved, and can we see the accuracy of that judgement?
What is the repeat contact rate for conversations you counted as resolved?
Can you separate escalations by reason?
What happens to the number if we make "talk to a human" easier to find?
The last one is the most revealing question on the list. A vendor whose headline number falls when customers can reach a person more easily has been selling you a routing outcome and calling it service.
Reference: the metrics that actually run a service operation
The four numbers above are the ones misused in sales conversations. This is a different list, and a reference rather than an argument: the measures a service team uses week to week. Most never appear in a vendor pitch because they are harder to make look impressive.
Metric | What it tells you | Watch for |
|---|---|---|
First contact resolution | Share solved in one interaction | Improves if you define "contact" narrowly |
Repeat contact rate | Customers coming back about the same issue | The single best check on a containment claim |
Average handling time | Time an agent spends per conversation | Falls when agents rush, not only when they improve |
First response time | Speed of the first human or AI reply | An instant automated acknowledgement is not a response |
Resolution time | End to end, from contact to solved | Distinguish working hours from elapsed hours |
CSAT | Satisfaction with a specific interaction | Averages hide the worst experiences |
NPS | Willingness to recommend | Measures the relationship, not the interaction |
Customer effort score | How hard the customer had to work | The most useful single number, and the least used |
Backlog and queue age | Work waiting, and how long it has waited | Age matters more than volume |
Agent occupancy | Share of time agents are handling work | High occupancy predicts attrition |
Escalation rate by reason | Why conversations reach a person | Useless as a single figure, see above |
Channel mix | Where contact arrives | Shifts faster than teams re-staff |
Contact rate per order or account | Contact volume relative to business size | Rises with growth; absolute volume misleads |
Reopen rate | Conversations closed then reopened | The honest counterweight to resolution |
Cost per contact | Total cost divided by volume | Falls when quality falls |
Deflection to self-service | Contacts avoided by content | Only meaningful with the repeat rate beside it |
Knowledge coverage | Share of questions your content answers | Rarely measured, quietly decisive |
Automation rate by intent | What is automated, broken down by type | Far more useful than one overall figure |
Pick fewer than you think. A team tracking six of these well is in better shape than one reporting all eighteen and acting on none.
Where to start
None of this means the numbers are worthless. Measured honestly and read alongside the repeat contact rate, they will tell you a great deal about where your service operation is working. It only goes wrong when a percentage is treated as a benchmark rather than a description of one company's query mix, denominator and definition.
So agree what you mean by resolved. Write it down. Publish it internally so everyone is reporting the same thing. Then compare vendors on those terms, and the conversation gets considerably shorter.
If you are comparing AI customer service vendors, the fastest way through the percentages is to agree the definitions first, then ask each vendor to report against them.
Frequently asked questions
What is the difference between deflection and resolution?
What is containment rate in customer service?
What is a good AI resolution rate?
How do vendors calculate AI resolution rates?
Which customer service metrics should I actually track?
Further reading

The best service conversation may be the one you start first
Which proactive messages earn their place, and where over-sending starts.
Read the guide

Customer service automation: what to automate, what to escalate, and what it costs to get wrong
How to decide which customer service work to automate, and which to escalate.
Read the guide

Empower Customer Service in 2026: The Strategy
Give your team one unified inbox, automate routine queries with AI Agents, and free agents for the high-value conversations. Here's the strategy, end to end.
Read the guide



