How do you evaluate a voice agent when a transcript-correct answer still fails the caller? At Bigcircle we treat a correct transcript as necessary and still incomplete. Turn-taking, interruptions, whether the speech could be followed, and how the silence felt all sit outside a text rubric. A call passes only when the text checks and the audio checks both pass, and a change has to be shown on the same recorded calls before it takes traffic.
The hard part is that the failure never shows up in the record a team scores. A judge reading the transcript sees the right amount and a polite close, and it marks the call passed. Nothing in a transcript-only trace can turn that pass into a fail, because the collision was never written down as a field the judge could read.
On the recording the caller waited through a silence, started a second sentence, and was talked over by the reply they had already begun to repeat. The words of the answer were fine and the turn was lost, and the caller is the person who decides whether to phone again about the same request. A green transcript score beside those repeat calls is the disagreement this guide is about.
The second difficulty is the test people trust before they trust production. A demo call is quiet, short and scripted, and the person on the other end waits their turn. Production callers talk over the agent, pause to find a number, and phone from cars and kitchens, so a prompt edit that wins on the demo can leave those calls untouched while the transcript scores stay green.
Listening to a handful of bad recordings will show you the shape of that gap, and it will not tell you whether the next build fixed those calls or a different handful. A hunch cannot be replayed. The comparison needs the same files, with the audio still attached, every time you try a change.
This guide starts with what a transcript judge cannot see. It then separates content failures from experience failures, and sets out the audio signals that place a call in one cell or the other. It covers how those signals sit next to text judges without being averaged, how a frozen set proves a fix, and what each voice architecture leaves you able to measure. How the agent is built is a different guide.
What do transcript judges miss?
Transcript judges miss the timing of the turn, the interruptions, and whether the caller could follow the speech. They read the words after the call and score whether those words were right. A caller who sat in silence, got talked over, or could not catch a number has already had a failed call, and the transcript can still look correct. The miss is in the audio, so the eval has to score the audio.
Take a claim-status call as the shape of the failure, rather than as a measured result from a line we publish. The caller asks where the claim stands, and the transcript shows a correct status, the next step and no policy breach. Anyone reading only that transcript would pass the call, and a rubric written with care would pass it too.
The timeline of the same call shows a gap after the caller stopped, long enough that they begin again, and then both sides are speaking. A text judge receives one question and one answer, and nothing in that input can mark the collision, because the timestamps never reached it. If you add a sentence to the rubric asking whether the call "felt natural", the judge still has no audio, so it will invent a feeling from the prose.
That blindness is why a green transcript run can survive a week of complaints. The complaints name waiting, being cut off, and not catching a reference number, while the eval names answer quality, and the two lists never meet. The team then edits the prompt, the transcript score rises, and the complaints stay, because the prompt was not the defect.
Confident AI described this gap in its Launch Week 03 post on voice evals, announced 21 September 2026. A transcript can show the right answer while the caller still cannot follow the call, because of slowness, a missed turn, noise, or speech that is hard to understand. We treat that as their product claim about the kind of signal a transcript lacks, not as a measurement of any line of ours.
The practical consequence sits in the trace you keep before you write another rubric. A transcript-only record cannot be rescued by a better prompt for the judge, however you phrase the question. You need the audio, the end-of-turn decision, and the timestamps of who was speaking, or the taxonomy in the next section has only one column and every experience failure is filed as a pass.
When is a call content-ok and still a bad experience?
A call is content-ok and experience-bad when the words answer the request and the caller still could not use the conversation. Content covers the facts, the policy and the tool result. Experience covers whose turn it was, how long the silence lasted, and whether the speech could be followed. Split the two, or a wording fix will take credit for a call the caller could not finish.
This figure is a framework. It names four outcomes of one call and reports no frequencies, because the mix depends on your line and we do not publish one.
| Cell | What the caller got | What a transcript judge says | What has to change |
|---|---|---|---|
| Content right, experience usable | A correct answer they could follow | Pass | The only cell a release may grow |
| Content right, experience unusable | The right words, on a turn they could not use | Pass, wrongly | An audio check has to fail this call |
| Content wrong, experience usable | A clear call with a wrong fact or a policy miss | Fail | A text rubric, one question in each judge |
| Both wrong | A wrong answer they also could not follow | Fail | Two defects, fixed on their own lines |
The cell that misleads a team is the second row. Every transcript dashboard is green there, and the usual response is to leave the audio path alone and tighten the prompt. The caller could not take a turn, or could not hear the number they were asked to confirm, and a more careful sentence does not give them the floor back.
If you ship that prompt change, the cell does not move, and the next week of complaints is evidence you scored the wrong column.
The fourth row has the opposite trap. Two checks fail, someone folds them into one number, and a large content miss drowns a smaller overlap. You then ship a factual fix, the blended number jumps, and the overlap is still in the recording. A call leaves the second row only when the audio verdict changes, and it leaves the third only when the text verdict changes. Count them as separate exits.
Our guide on how to build LLM evals covers how to find failure modes by reading traces before you freeze a rubric. Do that reading on the recording, not only on the transcript, or the open coding will never produce an experience category. The categories you skip at that step are the ones no later metric can invent, and a cell with no signal under it will stay empty while the calls keep failing.
Which audio signals should you score?
Score the audio on four signals: when the turn was judged over, how an interruption was handled, whether names and numbers were intelligible, and where silence and overlap sat inside the turn. Each signal points at a different stage. One voice-quality number mixes them, so a change that shortens silence can hide a worse interruption and you will not know which stage to open.
| Signal | Computed from | A failure means | It does not mean |
|---|---|---|---|
| End of turn | The detector's decision, set against the caller's next words | The agent started too early, or held on through a pause that was not an ending | That the answer text was wrong |
| Barge-in | Who spoke, whether playback stopped, and whether history was trimmed to what was heard | The agent talked through a real interruption, or it yielded to noise | That a tool call already sent was undone |
| Intelligibility | Word error against a reference, split out for names, numbers, dates and addresses | The words that drive the next action were not the words spoken or heard | That the whole conversation was the right one |
| Silence and overlap | Timestamps of each side, kept as a distribution per turn | A tail of dead air after the caller stopped, or overlap that was not a backchannel | That the mean latency of the day was fine |
End of turn is the decision that the caller has finished. A fixed silence gap works on quiet, fluent speech and fails when the caller pauses to read out a policy number, or switches language mid-clause, because those pauses are not endings. Score the decision against what the caller did next.
If they were still inside the same request, the detector cut in. If they had finished and the agent waited through a gap that made them start over, the detector held too long. Our guide on production voice-agent architecture covers how to build that detector. The eval requirement here is only that the decision is in the trace, so a later build can be compared on the same calls.
Barge-in is the caller speaking while the agent is still playing audio. Stopping playback is not the whole event. The history has to be trimmed to what the caller actually heard, and the interruption has to be classed as a backchannel or a new turn, or you cannot tell a polite "mm" from a caller trying to take the floor.
A hum that cuts the agent off is a false yield. A real question that the agent talks through is a missed yield. Score those as different failures, because the fix for one is the cause of the other.
An interruption also does not roll back a tool call that has already been sent, so a barge-in check that looks only at audio will pass a call where the refund already ran. Pair it with the tool log whenever the action is hard to undo.
Word error rate is the usual intelligibility proxy, and it is narrower than teams treat it. Use one reference for the caller's audio, which checks transcription, and a different comparison for the agent's speech: re-transcribe what was played and compare it with the text the agent meant to say, which checks whether the voice could be followed.
A single rate on the whole call hides a wrong surname inside a cloud of small function words. Score names, numbers, dates and addresses as their own slice, because those are the tokens that select the wrong tool call. The reference is only as good as the labelling, so audit a sample of the marked errors before you trust a movement in the rate. Some of them will be the label, not the model, and a "regression" that is really a relabel will send you into the wrong component.
Silence and overlap come from timestamps you can log on every turn, which is why they can run on the full traffic rather than a sample. Keep the distribution. A mean will look calm while a tail of turns leaves the caller in dead air, and those turns are the ones that become complaints.
We do not publish a silence cutoff for you to copy, because a gap your callers tolerate on one line is a hang-up on another, and the threshold has to come from the tail of your own turns. Overlap needs the backchannel class from the barge-in check, or every "mm-hm" looks like a collision and the metric trains you to ignore it.
These four can all pass while the answer is wrong, and they can all fail while the answer is right. Folding them into the text score erases that fact, which is why the next decision is which line owns which failure.
How do you keep audio scores from double-counting text judges?
Run text judges and audio checks on the same call, and store the two verdicts on separate lines. Text owns whether the answer, the policy and the tool arguments were right. Audio owns whether the caller could take the turn and follow the speech. An average counts one wrong fact more than once and lets a perfect answer hide a broken turn. The gate fails when either line fails.
This figure is a framework. It shows which verdict belongs to which check, and which team owns an audio failure. It is not a vendor's metric list, and it states no weights.
Double-counting is easy to build by accident, and it decides which bug you will chase. A judge asked whether the call was helpful will punish a wrong refund amount. A second judge asked whether the task completed will punish the same amount. Put a turn-taking score into that average and the factual miss dominates it, so the blended number will rise when you edit the answer, and the turns that collided will still collide.
Our evals guide already asks you to put one question in each rubric. Voice adds a second rule on top of that: do not ask the text judge about silence, overlap or intelligibility when a timestamp or a word-error slice already records them.
Online evaluation, in that same guide on how to build LLM evals, watches live traffic for behaviour the frozen set has not named yet. On a voice line the useful online signals are a caller who asks for a person, a call that ends in the silence tail, an overlap classed as a real interruption, and a repeat call about the same request. Use them to choose which recordings enter the set.
Asking an online LLM as judge to rate the "naturalness" of a call whose overlap and silence you already measured puts the silence in the score twice, and you still cannot tell which copy moved.
The two lines map onto the two operating layers in our guide to dual-layer production agent ops, and the mapping is not "audio means infrastructure". If the turn detector ran as it was configured and the call was still unusable, the caller did not get what they needed. That is a quality failure, owned by the team that builds the agent.
If a stage did not run, audio never arrived, or synthesis returned nothing, the system did not do what it was set to do. That is an infrastructure failure, owned by whoever pages on the path. One voice-quality alert sends both pages to one team, and that team will look at the layer it already knows.
The same split shows up in other agents we run. Our guide on building production-grade analytics agents scores whether the answer matches the owner's number separately from whether each gate fired, because a right-looking number can sit on a gate that failed open. A voice call needs the same pair of lines: the answer's truth, and the turn and the path.
Jev can sit beside them as a check on whether this caller should reach a person, run off the request path and in shadow first. Jev does not hear audio. It cannot be the turn-taking eval, and a high score must not override a caller who already asked for a person.
The deterministic timeline checks can run on every turn, because they are exact and they read logs you already have to keep for observability. A judgement of whether the answer was right stays on a sample, which the ops guide draws as the line between inline checks and later checks.
Sampling means a bad answer can reach a caller before anyone scores it, so anything that must never be said, such as a required disclosure or a tool argument, still belongs in code on the turn itself. The architecture guide puts those checks in the tool boundary. Here the requirement is that their pass or fail is a field on the call, not a sentence buried in a judge's reason.
How do you prove a voice fix on the same calls?
Prove a voice fix by replaying the candidate against a frozen set of recorded calls. The same audio, the same labels and the same traces go through every build you compare. The change holds when calls that failed an audio check now pass, calls that passed still pass, and the text verdicts are not worse. A new synthetic call can suggest the change. It cannot be the evidence that passes it.
This figure is a framework. It is the order we require before a voice change takes live traffic, and it reports no pass rate from any line.
A frozen call is a file you can replay: the caller audio, the agent audio when you have it, the transcript, the end-of-turn and barge-in decisions, the tool trace, and the label for which cell the call was in. Version that set the way you version the prompt. If someone edits a label after the change, the comparison measures the edit, and you will ship a relabel.
Redact personal data before a recording leaves the production store, which the ops guide covers, and keep the full recording under the retention rules you already apply to calls.
Replay makes two builds comparable. The candidate runs on those files, and you diff the verdicts per call rather than the average. An average can rise because a group of easy calls got slightly cleaner wording while the collision calls stayed broken, and the average will still look like an improvement.
Read the calls that flipped the way the evals guide describes for a sign test, and do not ship on a one-point move in a blended score. The direction has to show up on the calls you already labelled.
| Line | Passes when | Fails when |
|---|---|---|
| Audio failures | Each call labelled experience-bad on a named signal now passes that signal | Any labelled failure is still failing |
| Regressions | Calls that passed the audio checks still pass them | A passing call grows a new silence, overlap or cut-in |
| Text | Task, policy and tool verdicts are unchanged or better | An edit made an answer or a tool argument worse |
Synthetic callers are useful before you have the recording, because they can force an interruption or a noisy room you want to see once. They speak the way they were prompted, and your callers do not, so a simulation that improves is a reason to try the change on the frozen set. It is not a reason to widen traffic.
When a live signal names a failure the set lacks, add that recording and freeze the label before the next comparison. Adding it after you have seen the new build's output is how a set starts agreeing with the change you hope to ship.
Release stays a share of live calls, as the architecture guide describes, and the gate has to hold before that share grows. If a class of call, such as code-switched questions about a claim, starts failing after you widen, send that class back to people and leave the rest running.
The frozen set cannot contain a failure mode you have never recorded. The share limit is how you find out you missed one, without making it the whole line.
What does the voice architecture change about the eval?
The questions you ask of a call stay the same across architectures, and the fields you can score do not. A cascade gives you text at each stage. A speech-to-speech model gives you audio and a side transcript. A managed pipeline gives you whatever the vendor exports. Score the signals you can actually log, and do not treat a missing field as a pass.
Our guide to production voice-agent architecture compares a cascade, a speech-to-speech model and a managed pipeline on latency budget and on control of each stage, and it is the place to choose one. This section only says what that choice does to evaluation. Rebuilding the turn-taking design here would hide the visibility each shape leaves you when a call fails.
A cascade runs speech to text, a language model and text to speech as separate stages. Each boundary produces text, so you can fail transcription, the answer and the spoken audio independently.
Word error on the caller audio is a transcription check. The text judge reads the model output. A second intelligibility check on the played audio tells you whether the voice said what the model wrote. When a call fails, the stage is named, and the next edit has somewhere to go.
A speech-to-speech model hears and answers in audio, and the transcript is a side output rather than the text it reasoned over. A text judge on that transcript is checking a copy. It can still catch a wrong fact when the copy is faithful, and it can miss a failure that never made it into the copy cleanly.
Lean on the audio signals, and say in the report that the content verdict is on the side transcript. Treating that verdict as if you had inspected the model's input will clear calls you did not inspect.
A managed pipeline usually runs the audio path for you and returns what it chooses to expose. You can score tool arguments and policy on your side of the boundary, because those stay yours. Turn decisions, overlap and per-stage time exist only if the vendor exports them.
If the export lacks an end-of-turn decision, you do not have an end-of-turn eval, and a green tool-argument score must not be read as a green call. Export the recording and the timestamps into your own store at the end of the call. A dispute months later cannot depend on a vendor's log retention.
| You need to score | Cascade | Speech to speech | Managed pipeline |
|---|---|---|---|
| Answer and policy | Model text | Side transcript, labelled as a copy | Whatever text the vendor returns |
| Caller transcription | Your speech-to-text output | Only if a caller transcript is exposed | Only if it is exported |
| Agent intelligibility | Played audio against the text you sent | Played audio, with no separate text stage | Played audio, if you can record it |
| Turn-taking and overlap | Your detector's decision and timestamps | Model behaviour, from the recording | Vendor events, if they exist |
Whichever row is empty is a hole in the gate, not a silent pass. If you cannot fill it, keep the class of call that depends on it with people until you can, or pick an architecture that logs it. The architecture guide's release by share of calls still applies on top of that hole. An eval that cannot see the turn should not be the reason a share goes up.
Common questions
Why does a correct transcript still fail a voice agent?#
A transcript judge reads the words. The caller hears timing, interruptions, and whether a name or a number was clear. If the agent waited too long, talked over the caller, or spoke a number they could not catch, the words can still be right and the call has failed. Score those on the audio, and fail the call when either the words or the audio fail.
What is a content-ok, experience-bad call?#
The answer, read as text, is right and within policy, and the caller still could not use the call. Typical cases are a long silence after they stopped speaking, the agent talking over them, or speech they could not make out. Treat that as its own failure. A change that only improves the wording has not fixed it, and a transcript score will keep passing it until an audio check is allowed to fail the call.
Is word error rate enough to evaluate a voice agent?#
Word error rate compares a transcript with a reference and says how many words differ. It does not see turn-taking, silence, or whether the caller was interrupted. Use it as a check on transcription, and as a separate check on the agent's own speech, and score names, numbers and dates on their own. A transcript can be nearly all correct and still get the one identifier wrong, and that identifier drives the tool call.
How do you stop audio scores and text judges from double-counting?#
Give each check one question. Text judges own the answer, the policy and the tool arguments. Audio checks own the timeline and whether the speech could be followed. Keep the verdicts on separate lines and fail the gate when either line fails. A blended score lets a perfect answer hide a broken turn, and it counts one wrong fact twice if two judges both punish it.
What goes into a frozen voice eval set?#
Recorded calls from your own line, with the audio, the transcript, the turn decisions and the tool trace kept together, plus a label for which failure the call shows. Freeze the files and the labels, and replay every candidate through that same set. Add a call when production shows a new failure, and freeze it before you score the build you hope will fix it. Do not edit the set to make the change you just shipped look better.
How should a voice change be released?#
Replay it on the frozen set first. It passes when the calls that failed the audio check now pass, the calls that already passed still pass, and the text verdicts are not worse. Then widen the share of live calls, and send a class of call back to people if that class starts failing. A new synthetic call can suggest the change. It cannot pass it, because a prompted caller is not your caller.
Does a speech-to-speech agent need a different eval?#
The questions stay the same, and what you can log changes. A cascade gives you text at each stage, so you can score transcription, the answer and the spoken audio apart. A speech-to-speech model reasons over audio, and the transcript is a side copy, so a text judge is checking that copy. Score the audio either way. Say in the report when a content verdict was made on a side transcript rather than on text the model actually used.
Where do turn-taking failures sit in production ops?#
If the turn detector ran as it was configured and the call still felt wrong, that is a quality failure: the caller did not get a usable conversation, and the agent team owns the fix. If a stage did not run, the audio dropped, or synthesis returned nothing, that is an infrastructure failure, and platform owns it. Putting both in one voice-quality score sends the page to the wrong team.
Further reading
- Production voice-agent architecture. Cascade, speech-to-speech and managed pipelines, and the turn-taking design this guide does not repeat.
- Dual-layer production agent ops. Quality and infrastructure as separate owners, and how a failure becomes an eval case without leaking personal data.
- How to build LLM evals. Rubrics, judge checks, the frozen set and the sign test on flipped cases.
- Introducing Jev into agentic workflows. The shadow-then-gate sequence for an escalation check that does not hear audio.
- Building production-grade analytics agents. Scoring answer truth separately from whether each gate fired.
- AVOX, the insurance voice agent. A published voice line whose latency budget is not, by itself, an eval result.
- Voice agents, evals, LLM as judge and speech to text.
- Confident AI, Introducing Voice AI Evals: Test the conversation, not just the transcript, Launch Week 03, announced 21 September 2026. Their claim that a right transcript can still be an unusable call, and that deterministic audio metrics can sit beside text judges.