Skip to content
fiveleaf
Answers/

Measurement

·

2 min read

How do I know if my AI agent is working?

Silviu Major·Founder, Fiveleaf·

Containment rate on its own tells you almost nothing, and it is the number most vendors lead with.

Containment measures that a conversation was not escalated to a human. It does not measure whether the customer got what they needed. Someone who gives up in frustration and closes the window counts as contained, and so does someone whose problem was solved perfectly. The number cannot tell them apart.

The three you need together

Containment. Did the agent handle it without a human.

Satisfaction. Did the customer feel served. If containment rises while CSAT falls, you are not automating, you are deflecting, and the trend line is telling you so.

Recontact rate. Did they come back with the same issue within a day or three. This is the clearest evidence that a contained conversation was not actually resolved, whatever the first number said.

Read as a set they are meaningful. Read alone, containment is a vanity metric and everyone in the industry knows it.

Benchmarks worth holding against

Most chatbots contain 20 to 40%. Mature well-integrated agents reach 70 to 90%. Industry-average CSAT for AI support sits around 78%, with leaders above 85%, roughly level with live chat.

Median tier-one deflection lands nearer 41% than the numbers on vendor slides, with the top quartile around 59%.

And aggregate figures mislead badly, because simple intents like password resets deflect above 70% while nuanced complaints rarely break 25%. A headline rate without the query mix underneath is close to meaningless.

What I would actually watch in the first month

Not the dashboard. The transcripts.

Read fifty conversations a week, properly. You will find the failures before any aggregate does, and they will be specific and fixable: a question phrased in a way the agent did not expect, a policy it does not know about, a hedge at the end of a correct answer that is pushing people to ask for a human.

We had exactly that one. Resolution sat flat for weeks while transcripts looked fine, because the agent was adding "please confirm with our team" to answers that were already right.

You do not find that in a chart.

The commercial number

Eventually you want this in money, and the honest version is smaller than the marketing version.

Apply the benefit only to the contact the AI actually handled, not to total volume. Use gross margin rather than revenue for any sales uplift. Subtract the running costs, meaning platform, model usage, maintenance and integration work. Do not count the same contact twice across two different savings categories.

What you get is a figure you can defend to a CFO. Usually still good. Never as good as the slide.

Frequently asked

What is a good containment rate?
Most chatbots contain 20 to 40% of conversations and mature well-integrated implementations reach 70 to 90%. But containment alone only tells you the customer did not escalate, not that their problem was solved, so a number without CSAT and recontact beside it is unreadable.
What is the difference between containment and resolution?
Containment means the conversation was not passed to a human. Resolution means the customer's problem was actually fixed. A customer who gives up and closes the window counts as contained. Conflating the two is the most common measurement error in the category.
How soon should I expect the numbers to look good?
Not immediately, and be suspicious of a deployment that looks perfect in week one. Expect a few weeks of tuning after go-live where you are watching transcripts rather than dashboards, because the early wins and the early failures are both visible in the conversations well before they show up in aggregate.

If you want help building this

Building AI agents into a mid-market business is what Fiveleaf does.

Bespoke build, fully integrated, continuously optimised. A 30-minute discovery call is enough to tell you honestly whether AI agents fit your team right now, or whether you’re better off waiting six months. No pitch.

About the author

Silviu Major, Founder, Fiveleaf

Silviu Major

Founder, Fiveleaf

10+ years building automation systems inside enterprise SaaS, now applying that same operational rigour to AI implementation for mid-market businesses. Writes about what works (and what doesn’t) from inside live deployments, not from the outside looking in.

Connect on LinkedIn →

Keep reading