Published

Which support metrics stop working once AI answers

First response time drops to zero and stops meaning anything. Deflection rate can be raised by hiding the contact button. What to measure when half the conversations are read by nobody.

Adrià Castany 10 min

Ilustración de barras crecientes en los colores de Intake

The first month with an automated agent, the dashboard looks beautiful. First response time falls from four hours to eleven seconds. The volume reaching people halves. Deflection rate comes out at 40%.

And then somebody asks whether support is actually going better, and it turns out those numbers cannot answer that.

It is not that the metrics lie. It is that most of them were designed to measure a team of people with limited capacity, and the moment part of the conversations stop passing through a person, they stop measuring what they measured.

The ones that break

First response time

It measured the wait until somebody had read you. With an agent answering instantly, it goes to zero and stays there forever.

The number keeps appearing, which is what makes it dangerous: it looks like a metric and it is not one any more. A dashboard where first response time reads eleven seconds every month tells you nothing, neither when support is going well nor when it is going badly.

What to measure instead: the time to the first useful response. That is, the response the customer does not have to rephrase. If it answers instantly and the customer writes the same question again in different words, that first response did not happen.

And separately, the one that matters most: the time to reach a person when a person is needed. That is where the number can degrade unnoticed, because the cases needing somebody are a minority and they hide behind the average.

Deflection rate

It is the metric used to justify the investment and the easiest one to inflate without meaning to.

The problem is not the definition, it is the denominator. If you count as deflected any conversation that never reached a person, you are also counting whoever gave up. And hiding the contact button raises deflection rate and damages the business, in that order.

How to measure it honestly: conversations closed without human intervention and without the customer returning within seven days, over the total conversations started. The second condition is what separates the metric from the shop-window version.

The check that confirms it in a minute: if deflection rate rises and CSAT falls the same month, you have not deflected, you have hidden. Same if it rises alongside the number of customers writing through another channel — the account executive's email, LinkedIn, the sales chat.

CSAT

It still works, but it measures something else and has to be read differently.

When a person answers, CSAT measures the whole interaction: the tone, the judgement, the resolution. When an agent answers, it splits into two cases that should not be averaged:

  • It resolved: the score usually rises, because the answer came sooner.
  • It did not resolve: the score falls further than it would have with nothing there, because the customer experiences the failed attempt as an obstacle placed on purpose.

Averaging the two gives a middle number describing neither. Segmenting by who resolved — agent, person, agent then person — is the minimum for CSAT to keep saying something.

Tickets per agent

It stops meaning workload. If the automated agent takes the repeated and documented work, what reaches people is harder by definition: the cases needing judgement, the angry ones, the ones nobody had seen.

A team handling half the tickets it did six months ago may be working more, not less. Measuring volume per person without looking at the composition of what is left leads to the opposite of the correct conclusion.

The ones that start mattering

Real resolution rate

Of all the conversations the agent answered, how many were effectively closed: the customer did not write again, did not rephrase, and did not ask for a person.

It is the central metric of an AI system and one that almost never ships in dashboards by default, because it requires looking at the following days rather than the moment of closing.

Escalation quality

When the agent hands a conversation over, does the case arrive assembled or blank?

It is measured with one concrete figure: how many escalated conversations force the customer to repeat something they already said. It is the metric that best separates an escalation that saves time from one that creates it, and it explains why two companies with the same deflection rate have very different satisfaction.

Rephrasing rate

How often the customer writes the same question again in different words after an agent response. It is the earliest signal that the answer did not land, and it shows up weeks before CSAT moves.

It is also a gift to the documentation: the list of rephrasings, grouped, is literally the list of articles that are badly written or missing.

Cost per resolved conversation

The figure that settles the argument, and the only one comparable between before and after.

It has two terms: what the agent costs — subscription plus usage — and what the human time still required costs, using real handling time including the context hunt and the interruption.

Without this metric, the conversation about whether AI pays off gets decided on impressions.

Coverage by question type

Not a global percentage, but the breakdown: of the billing questions, how many resolve on their own; of the configuration ones, how many; of the integration ones, how many.

It is what turns the metric into a decision. A global 30% does not tell you what to do; "80% of billing questions resolve and 10% of integration questions do" tells you exactly where next week's work is.

How to build the dashboard

Four numbers, not twelve:

MetricWhat it answers
Real resolution rateIs it working?
Cost per resolved conversationDoes it pay off?
CSAT segmented by who resolvedDoes it help the customer?
Coverage by question typeWhere is the next piece of work?

And one warning worth more than the four: none of this is comparable with last month if the question mix has changed. A launch, a price rise or a campaign bringing in customers from another profile change the composition of what arrives, and with it every average.

Before switching anything on

The only way to know whether something improved is having the "before". And the before has to be measured by hand, because it is not in any dashboard.

A hundred real conversations, sorted into three piles: the ones answered with general information, the ones depending on account state, and the ones needing judgement. That split sets the real ceiling of what can be automated, and it is what turns any target into something reachable or into a promise somebody will break in three months.

With that split in front of you, the discussion stops being about which tool and becomes about which part of the work is coming off your plate. Which is the only version of the conversation that goes anywhere.