Field Notes · Aventary Exclusive

The Metrics Became the Product BacklogWhat a conversational AI system taught me about measuring RevOps.

I put a conversational AI agent into production. It answers phone calls and it messages people inside the app. The first question everyone asked was how many calls it handled. That turned out to be one of the least useful things I could measure.

Call volume tells you the thing is running. It does not tell you the operation underneath it is getting better. Those are two different questions, and the gap between them is the whole subject of this piece.

So I threw out the dashboard I started with and rebuilt it around the machinery. Here is what I found, what it cost me to admit, and why any RevOps leader should care.

A clean dashboard can sit on top of a broken handoff for months.
01 / The Wrong Dashboard

The one I started with measured the wrong things.

Every one of these can look healthy while the operation quietly falls apart. The agent can answer every call, keep each one short, and hand it off politely, and the person on the other end is still repeating themselves, calling back four minutes later, and getting bounced around with no context.

Call volumeContainmentAvg durationTransfer countSentiment

None of the vanity metrics could see that. So I replaced them with five that describe how the work actually moves.

02 / The Machinery

Five metrics that describe the machinery.

  1. 01
    Caller repeat rate. After a transfer, did the person have to re-explain the whole thing from the top?
  2. 02
    Context handed off. Did the human on the receiving end actually get the story, or just a warm body on the line?
  3. 03
    Five-minute callback. Did the same person call right back because nothing got solved?
  4. 04
    Longest calls, by type. Which request types eat the most human time every week? That is your roadmap hiding in plain sight.
  5. 05
    Confirmed human answer. A transfer attempt is not proof a human answered. Measuring the attempt as a win is grading yourself on effort.

Now the receipts, and I want to be honest about how thin they are.

39→53%
Callers who clearly had to repeat the issue after transfer. It went up, not down, and that is the honest read.
12–15%
Context that clearly survived the handoff to the human.
11–27%
Same-person callbacks inside five minutes, depending on the week.

Two-week transcript audit. 67 audited calls, then 58. A small sample, and I am not going to pretend it is the whole system.

Every ugly number pointed at something specific to build.
03 / From Metric to Build

The metrics stopped being a report and became a build list.

I read the longest, ugliest calls and asked one question on each. What should have made this call unnecessary? The answers were not strategy. They were a backlog.

Repeated time-off calls

A person re-collected name, dates, and schedule from scratch on every call.

Built A structured leave-intake that captures all of it before any handoff.

Long supply calls

The agent understood the request but could not complete it, so it sent a human.

Built A supply request written straight into the system of record.

"The app is not showing my info" calls

A person diagnosed it live, on every call.

Built A status diagnostic and an escalation packet built into the flow.

"Account status" calls

One vague label was hiding payroll, training, and supply questions.

Built Four real intents, each with its own routing.

04 / The Trap in the Average

The number that went up for a good reason.

Average call duration climbed, and the reflex read is that things got worse. When I segmented the calls, the opposite was happening. The AI was resolving the short, easy calls and closing them out, so those stopped reaching a human at all. What was left in the queue were the genuinely hard ones. The average rose because the easy volume left, not because any single call got worse.

The reflex read

Duration is climbing. The system is getting worse.

What was actually happening

The AI cleared the easy calls, so only the hard ones were left. The mix changed, not the calls.

That reframe changed the work. Instead of chasing the average down, I let it climb and went after the hard problems it exposed, one pattern at a time.

Repeat callers told the same kind of story. Watching who called back more than once, some of it was not a phone problem at all. It was adoption. So those callers got redirected to in-app messaging, which cut the repeats and resolved the need in a channel that fit it better.

Sentiment earned its place for a reason most dashboards miss. It measured more than how a caller felt. It told me whether a change I had just shipped was landing, whether people were adopting the new tooling or working around it. Self-addressed by AI is the metric I optimize for. Sentiment is how I know whether the work behind that number is holding.

05 / The RevOps Bridge

The same trap lives in your funnel.

Most RevOps dashboards are built like my first one. They report what happened at the edges. Leads, meetings, pipeline, conversion, forecast. All real, all lagging, and all silent about the machinery that produced them.

The friction lives in the middle, and the middle is finally measurable. AI can read the messy handoffs that used to be invisible. Most teams are using it to populate the same dashboard they built ten years ago, only faster. Point the same lens at your funnel instead.

Run these against your own stack
  1. How often does a lead get reassigned before the right owner actually engages?
  2. How much context dies between the SDR and the AE on handoff?
  3. How many opportunities move backward a stage after someone called them qualified?
  4. How often are reps re-entering data the company already has somewhere?
  5. Which campaigns create the most rework, not just the most leads?
  6. Which support tickets are expansion or churn signals that nobody ever connected to revenue?
  7. Which CRM fields are technically full and completely untrusted by the people using them?

Every one of those is a machinery question. Answer them with a number and you stop reporting last quarter and start seeing what to build next.

06 / The Principle

What a dashboard is for.

A report tells you what already happened. A build list tells you what to do about it. The whole move here was turning one into the other, and it did not take a new tool. It took measuring the operation instead of the technology.

A modern operating dashboard should not just explain last quarter. It should tell you the next thing to build.

Find the friction your dashboard is hiding.

The Aventary AI Diagnostic is short and fixed-scope. I read your operation the way I read those calls, show you where work is leaking between systems and handoffs, and hand you a build list you can act on. You get it in weeks.

Book Your AI Diagnostic
AVENTARY INSIGHTS · Field Notes · One conversational AI system, read as a build list.
The Metrics Became the Product Backlog — Aventary