Key takeaways
- Across two live calls, the agent’s answers were rarely the problem. Every failure we heard came from the voice pipeline around the model: detecting speech, deciding when the caller had finished, and handling interruptions.
- Once the caller’s final transcript arrived, most turns produced their first outgoing audio within 0.5 to 1.4 seconds. A few took 2 to 4 seconds, and of those only one came down to the model.
- Every call opens with an AI disclosure and a recording notice, and every write tool waits for the caller’s yes. Each call ends with a signed result a FlowX.AI process can act on.
- Structured output is not validated output. A score the schema bounds to 0 to 10 came back as 45 in JSON of the right shape, so the platform now checks values against the schema before it hands them on.
We expected the hard part of a voice agent to be the conversation. It turned out to be the pauses between turns.
We connected an Agent Builder agent to a phone line and had it place real calls to a mobile phone. It opened calls in different languages, switched when the caller asked, and read back names and amounts. At the end of each call it filled in a result form. The replies were short and on topic, and the model usually returned its first token in under half a second.
In these calls, the failures we heard were all about timing. The bot started talking while the caller was still speaking and got cut off after 160 milliseconds. Once it stayed silent for five seconds after the caller paused: it had heard sound but no words, and nothing closed a turn without words until a backstop timer ran out.
By turn-taking we mean everything at the edge of a turn: hearing speech, deciding it has ended, waiting for the final transcript, starting the reply, and yielding when the caller cuts in. This is a field report, not a benchmark: two live calls with one tester, plus scripted evaluations. It follows our browser agent benchmark and applies the same rule. Report what we observed, including the parts that went wrong.
1. What we built
A voice agent in Agent Builder is a Voice Agent node in an ordinary workflow. The node holds the system prompt, the tools, the languages, and the result schema. A separate voice service handles the audio. It is built on an open-source framework for real-time voice pipelines.
One turn goes through five stages:
- Telephony. The phone line streams the call’s audio to the voice service in both directions.
- Voice activity detection. A small model decides, frame by frame, whether someone is speaking.
- Speech to text. The caller is transcribed while they speak.
- The agent. The same agent Agent Builder runs in any workflow, with the node’s tools, knowledge and model.
- Text to speech. The reply becomes audio one sentence at a time, while the model is still writing the rest.
Between speech to text and the agent sits the part this article is about: turn detection. It decides when the caller has finished and the agent may answer. We use two strategies. Where a smart-turn model supports the language, it judges from the audio whether the utterance sounds complete. Elsewhere, a fixed silence timeout ends the turn.
2. Where the time goes
We log four timestamps for every turn: the caller’s final transcript, the start of the model’s reply, the first synthesized audio, and the moment the bot starts speaking. The figure lays out a typical turn from the second call, which used the silence timeout.
The per-turn number we report starts at the final transcript and stops when the bot’s first audio leaves our server. It leaves out the silence check and the transcript finalizing before it, and the phone network after it. So a caller waits at least 0.6 to 1.3 seconds longer than the table shows, before any phone network time. We did not measure the full delay from the caller’s mouth to their ear.
The silence check and the turn decision are two separate waits. The first happens before the transcript: the voice detector waits for a stretch of silence. The second happens after it: the turn detector decides the caller is done. With smart-turn, the second wait mostly disappears, because the decision is made while the transcript is still arriving.
| Call | Turn detection | Turns | Per-turn times |
|---|---|---|---|
| 1 · language switch mid-call · 104 s | smart-turn | 9 | 0.54, 0.55, 0.58, 0.69, 1.33, 2.17, 2.36, 2.80, 3.94 s |
| 2 · language switch mid-call · 88 s | silence timeout | 6 | 1.02, 1.10, 1.18, 1.31, 1.36, 2.84 s |
In these calls the model was rarely the slow part: it usually returned its first token in under half a second, and speech synthesis added about a tenth of a second. Even so, no turn came in under 0.5 seconds on the per-turn clock.
The two calls differ mainly in turn detection. Smart-turn decides in parallel with transcription, so when it was confident the reply started about 15 ms after the transcript. That produced the four turns under 0.7 seconds. The silence timeout waits about half a second after the transcript on every turn. It was slower and more even.
The slow turns had ordinary causes. The 3.94-second turn was a reply queued behind the opening line, the 2.84-second turn was a language switch, and the 2.17- and 2.80-second turns were turn detection waiting for the caller to go on. Only one, the 2.36-second turn, was the model being slow.
The turn-detection waits need a closer look. Of the two waits of about 1.7 seconds in the first call, one came before the language switch and one after it. Turn detection kept the smart-turn model chosen for the opening language, and the new language was one that model was not trained on, so part of that waiting may have come from running smart-turn outside its languages. Turn detection should change with the language. That is the next fix. It will not remove every long wait: the other 1.7-second wait came before the switch, in the opening language.
The longest wait has no row in the table, and the logs show why. The caller made about two seconds of sound that speech recognition turned into no words. Smart-turn reached its 3-second limit, but it ends a turn only once a transcript has arrived, and none did. So this was not smart-turn misjudging the pause. A separate backstop closed the turn at 5.0 seconds. With no words there was nothing to answer, so no reply started and no reply time was recorded. The caller spoke again a third of a second later.
3. Two ways to lose the floor
Timing problems don’t show up in a transcript. Reading one, both calls look fine. Listening to the first, the tester described it as “a bit choppy, a bit of a wait.” The logs showed two separate faults.
A third, thirteen seconds in which the bot said nothing, was most likely background noise during the test: the voice detector kept reporting speech while speech recognition produced no new words. It wasn’t a turn-taking fault between caller and bot, so we cover it separately. It still exposed a gap: turn detection relied on the voice detector alone. There is now a fallback in production: if speech recognition produces no new words for 2 seconds, the turn ends and the agent answers what it heard. It has not yet run on a live call with background noise.
Cut off after 160 milliseconds
The caller said “Yes.” The model took 2.2 seconds to start its reply: the 2.36-second turn in the table above. By then the caller had kept talking: “Please… tell me this agent,” in the transcriber’s words. Barge-in is a feature: callers must be able to interrupt a long reply. But the bot was stopping for any sound at all. One word, 160 ms into the reply, ended it. The bot then restarted with a different sentence, and to the caller the call sounded choppy.
While the bot is talking, it now takes two or more transcribed words to interrupt it. A stray “please,” a “mhm,” or echo on the line no longer counts. In the second call, the caller’s “Alright.” during a reply was logged and ignored, and the reply finished. The one exception is a yes or no to a pending confirmation question, which always counts.
The rule has a cost. A caller who says only “Stop” or “Wait” while the bot is talking can no longer interrupt it. The bot finishes its sentence first. We accepted that trade for now, and the word count is a setting.
A 5-second wait
Smart-turn judged some of the caller’s pauses as unfinished sentences and kept waiting for more: about 1.7 seconds on two turns, against a limit of 3 seconds, the framework’s default. The longest silence had a different cause: the caller heard five seconds of nothing because the turn had no words in it, as the previous section shows. On a phone call a gap that long feels like the line has dropped. We lowered the limit to 1.5 seconds, half the default, and the backstop from 5 to 2.5 seconds.
Both values are judgment calls, not tuned results. They are longer than the ordinary pause between sentences, and short enough that the gap doesn’t feel like a dropped line. The cost falls on slow speakers: a caller who pauses longer than 1.5 seconds mid-sentence may be answered before they finish. The 2.5-second backstop adds no cost of its own when the caller’s words are recognized, because the limit ends that turn first. The limit is a setting, so a use case with slower or older callers can raise it again. The backstop is fixed for now.
Three timers now decide when a turn is over, and each one catches a different failure:
| Timer | Watches | Fires after | Limits |
|---|---|---|---|
| Smart-turn limit | silence after a pause smart-turn judged unfinished | 1.5 s | ends the turn only once a transcript has arrived; languages without smart-turn use a fixed silence rule instead |
| Backstop | no voice-detector event and no transcript | 2.5 s | ends even a turn with no words, but never while the voice detector reports speech |
| No-new-words fallback | no new words from speech recognition in an open turn | 2 s | ends the turn even while the voice detector reports speech, and answers the words heard |
On an ordinary pause with words in hand, the smart-turn limit or the silence rule fires first. The backstop catches a turn with no words, like the five-second silence, and now fires at 2.5 seconds. The fallback catches a voice detector stuck on noise, like the thirteen-second silence. The backstop never fired there, because the detector kept reporting speech the whole time.
What the fix does not change: a turn with no words still gets no reply. The backstop closes it sooner, but the bot then says nothing, and the caller hears silence until they speak again. Asking the caller to repeat is the obvious next step, and we have not built it yet.
4. Switching language mid-call
Each node lists the languages a call may use, starting with the one it opens in. When there is more than one, the agent gets a switch_language tool. Calling it retunes speech recognition and speech synthesis for the new language, and from then on the agent replies only in that language.
Both calls used it. Each time the caller asked for another language partway through, and speech recognition, still tuned to the first one, garbled the request. The agent understood it anyway and switched.
Two smaller problems came out of this:
- A filler in the wrong language. The agent normally says “One moment, let me check” while a tool runs, and it said so in the old language just before switching. Call-control tools (language switch, stop recording, transfer) are instant, so they no longer get a filler.
- Too informal. In a language with a polite form of address, the agent used the familiar one. The prompt now asks for the polite form in every language, unless the caller asks otherwise.
Not every voice speaks every language, so a call only offers to switch to languages its configured voice supports.
5. Rules every call follows
Some behavior should not depend on how well a prompt was written. The node adds these rules to every call, and the platform enforces the ones it can without relying on the model.
We checked these rules with 24 scripted conversations, run as text only. They exercised the agent and its tools, not the voice path, so speech recognition errors, barge-in and turn detection played no part. 21 of the 24 passed. In all 24, no write tool ran without the caller’s yes. Read-back held less well: two of the three failed runs skipped it. In one, the agent did not read back the time and name, and never made the booking the caller asked for. In another, it never ran the order lookup it needed. In the third, it asked to confirm a cancellation without reading the booking back first, and offered a new appointment nobody had asked for. The confirmation is enforced by the platform; read-back is an instruction to the model, so it can be skipped.
An earlier batch of the same runs, before the fix below and not counted in the 24, found a bug of ours. The proxy between the agent and the LLM call dropped each tool’s list of required fields, so nothing made the agent ask for the caller’s name before saving a lead. In two runs it saved the lead with no name. In another it filled the name in with “Nepoznato,” Croatian for “unknown.” The required fields now reach the model.
Disclosure first. The opening line says the caller is talking to an AI assistant before anything else, in the call’s language. This addresses the Article 50(1) disclosure obligation of the EU AI Act. The agent is also told never to claim to be human.
Recording notice with a real opt-out. When a call is recorded, the opening says so and asks whether that is all right. If the caller objects, then or later, the agent calls a stop_recording tool. No audio is uploaded for that call, and the result says the caller declined.
A spoken confirmation before any write. Tools not marked read-only, such as saving a lead or booking something, pause the agent. It asks the caller to confirm, and the tool runs only on a yes. A yes that adds details, such as “Yes, but my name is Ana,” sends the agent back to redo the action with the correction. It does not run the stale version.
6. Handing the call to a process
A voice call is rarely the whole job. Someone has to book the meeting, send the information, or add the number to a do-not-call list. In our setup, a FlowX.AI process does that work. The voice agent isn’t the workflow. It’s a conversational step inside the workflow. The call collects information, clarifies what the caller wants, and asks before acting. The process decides what happens next and does the writing.
The node declares a result schema, the fields the process needs. When the call ends, a model fills that schema from the transcript. The platform then posts the result, with the summary, transcript, recording link, and any context the process passed in, to a callback URL. The request is signed with HMAC-SHA256 over the timestamp and body. Because the timestamp is signed, a receiver can reject stale requests. Our reference verification code refuses anything more than five minutes old, which blocks a captured request from being replayed later. That check lives in the receiver, so every receiver has to make it. Delivery is attempted up to three times when the receiver times out, is rate-limited, or returns a server error, and the outcome is stored with the call. A retry can deliver the same result twice, if the receiver handled the first attempt but answered too late. Every attempt carries the same session_id, so a receiver should treat a second delivery for a session it has already handled as a duplicate. This is the delivery from the second call:
POST /voice-call-ended
X-Voice-Event: voice.call.ended
X-Voice-Timestamp: <unix seconds>
X-Voice-Signature: sha256=<HMAC-SHA256(secret, "<timestamp>.<body>")>
{
"event": "voice.call.ended",
"session_id": "<uuid>",
"direction": "outbound",
"status": "completed",
"end_reason": "hangup",
"duration_seconds": 88,
"recording_declined": false,
"result": {
"right_person": true,
"experience": "none",
"budget_eur": 0,
"preferred_language": "English",
"outcome": "send_info",
"score": 45
},
"context": { "process_instance": "test-p-1", "lead_id": "L-TEST" }
}Look at score. The schema asks for an integer from 0 to 10, and the model returned 45. The JSON had the right shape, but the value broke the schema’s maximum, and in our setup nothing checked it after generation. Structured output is not the same thing as validated output. A process that routes leads by score would have treated a poor lead as an excellent one. The extraction prompt now states the bounds, and any value outside the schema is returned as empty. A missing value is safer than a wrong one.
From there, it’s an ordinary process. The context field carries the process instance, so the result lands on the right case. In the process we designed around these calls, a do-not-call outcome updates a list in the process’s own database, and a meeting request becomes a calendar booking through Integration Designer. The call itself never writes to either system. That process is designed but not yet connected.
7. What we took from it
In these calls the model was rarely the weak part. What decided whether a call felt right was the system around it: hearing speech, knowing when the caller was done, yielding when they cut in, confirming before acting, and checking the result before a process used it. Most of the engineering sits in the timing and the checks around the model. The model can know exactly what to say, and the call can still feel broken if the system doesn’t know when to say it.
Glossary
A few terms from voice pipelines used in this article.
| Term | What it means here |
|---|---|
| Turn | One stretch of speech by one side, ended by the other side starting to answer. |
| Turn detection | Deciding when the caller has finished, so the agent may answer. Most of the faults in this article came from it. |
| VAD | Voice activity detection: a small model that reports, many times a second, whether it hears speech. It can’t tell a finished sentence from a pause, and background sound can trigger it. |
| Smart-turn | An open audio model that judges whether an utterance sounds complete. It is trained on a fixed set of languages; for others we use a silence timeout. |
| Silence timeout | Ending the turn after a fixed length of silence. Simple and predictable, but it adds that silence to every turn. |
| Barge-in | The caller interrupting the agent mid-reply. Useful, as long as a single stray sound doesn’t trigger it. |
| Backchannel | Short sounds a listener makes without taking the floor: “mhm,” “yeah,” “alright.” |
| STT / TTS | Speech to text and text to speech. Here, both stream: text arrives while the caller speaks, and audio starts before the whole reply is written. |
| Filler | A short phrase the agent says while a tool runs, so the caller isn’t left in silence. |
| Result schema | The fields a node wants from each call, filled from the transcript when the call ends. |
| HMAC signature | A hash of the timestamp and body made with a shared secret. The receiving process recomputes it to confirm the result came from the platform and was not changed. |