Containment rate on its own tells you almost nothing, and it is the number most vendors lead with.
Containment measures that a conversation was not escalated to a human. It does not measure whether the customer got what they needed. Someone who gives up in frustration and closes the window counts as contained, and so does someone whose problem was solved perfectly. The number cannot tell them apart.
The three you need together
Containment. Did the agent handle it without a human.
Satisfaction. Did the customer feel served. If containment rises while CSAT falls, you are not automating, you are deflecting, and the trend line is telling you so.
Recontact rate. Did they come back with the same issue within a day or three. This is the clearest evidence that a contained conversation was not actually resolved, whatever the first number said.
Read as a set they are meaningful. Read alone, containment is a vanity metric and everyone in the industry knows it.
Benchmarks worth holding against
Most chatbots contain 20 to 40%. Mature well-integrated agents reach 70 to 90%. Industry-average CSAT for AI support sits around 78%, with leaders above 85%, roughly level with live chat.
Median tier-one deflection lands nearer 41% than the numbers on vendor slides, with the top quartile around 59%.
And aggregate figures mislead badly, because simple intents like password resets deflect above 70% while nuanced complaints rarely break 25%. A headline rate without the query mix underneath is close to meaningless.
What I would actually watch in the first month
Not the dashboard. The transcripts.
Read fifty conversations a week, properly. You will find the failures before any aggregate does, and they will be specific and fixable: a question phrased in a way the agent did not expect, a policy it does not know about, a hedge at the end of a correct answer that is pushing people to ask for a human.
We had exactly that one. Resolution sat flat for weeks while transcripts looked fine, because the agent was adding "please confirm with our team" to answers that were already right.
You do not find that in a chart.
The commercial number
Eventually you want this in money, and the honest version is smaller than the marketing version.
Apply the benefit only to the contact the AI actually handled, not to total volume. Use gross margin rather than revenue for any sales uplift. Subtract the running costs, meaning platform, model usage, maintenance and integration work. Do not count the same contact twice across two different savings categories.
What you get is a figure you can defend to a CFO. Usually still good. Never as good as the slide.
