Voice Agents · Q1 2026
This report summarizes the first full quarter of Voice Agents in production at commercial scale. Three cohorts grew, three new ones launched, and the operational signal across all six converged sooner than projected. What follows is what moved, why, and where we are heading into Q2.
Six production cohorts handled a combined 312,400 customer calls in Q1 2026. Voice Agents resolved 48% of those calls end-to-end — transactional steps included — without escalation. The remaining 52% followed the first-class human-handoff path, every escalation logged and reviewable. Operationally, this quarter is the first in which the Voice Agents product performs above the threshold required to redirect new customer engineering effort from onboarding onto cohort expansion1.
Three signals drove the quarter: resolution-rate stabilization across cohorts, a narrowing variance band between the best- and worst-performing deployments, and a sharper handoff profile — when the agent does escalate, it escalates faster and with better context attached to the ticket.
Resolution = call closed without human agent joining the line. Cohorts A–C entered Q1 already in production; D–F launched during the quarter.
Across the quarter, weekly audit volume climbed from 19,100 calls in the first full week to 27,800 in the last. That's a function of two things: cohorts D–F coming online in late January, and the early cohorts raising the cap on calls they route through the agent. The narrower story is about variance: the spread between the highest- and lowest-resolution cohort halved by quarter-end.
Volume is weekly call count routed through Voice Agents before any handoff decision. Excludes aborted lines and wrong-number filters.
The playbook convergence matters more than the headline resolution rate. A high-performing cohort without a reproducible rollout is a demo; a portfolio of cohorts landing within ten points of the median is a product. Q1 is the first quarter we've had the second.
Three priorities carry forward into Q2. Each is tracked against a numerical gate; none is marketed.
| Priority | Gate | Current | Target |
|---|---|---|---|
| Expand Cohort B & D envelopes | Call types accepted per cohort | 11 | 16 |
| Median escalation latency | Seconds from decision to human join | 14.2 | ≤ 10.0 |
| Evidence-per-escalation | Structured context fields attached | 3.1 | 5.0 |
| Rubric version drift | Weeks since last rubric publish | 4 | ≤ 6 |
It is easy to read the 48% resolution number as the headline. The more useful number is the 52% — the calls where the agent elected to escalate. In Q1, those escalations averaged 3.1 structured context fields attached to the handoff ticket; when an operator picked up the line, the customer was already identified, intent-tagged, and routed to the right specialist queue. That context is not incidental. It is the product argument.
In Q2 we push that number to five. Which sounds incremental until you watch a human agent pick up a Voice Agents handoff and read context instead of asking questions. The escalation stops being an interruption to the customer; it becomes a continuation of a conversation the AI already started well.
It is not a signal to expand headcount. It is not a reason to stop measuring. It is not evidence that the rollout playbook is complete — three cohorts (A–C) are an old sample; three (D–F) have been in production for less than 90 days. Statistical significance across the cohort set is in late-Q2 territory at the earliest.
Q2 is the first quarter where we ask the question Audibot is designed to answer: do the calls we didn't resolve look materially different from the calls we did?2 If they do, the rollout playbook extends. If they don't, we have a limit we need to explain.
SPEC-AUDIBOT-v2 §3.