pace- the ceiling on a turnendpointing_delay- the floor on a turn- Reading what you set - the resolved numbers
- How to tune it - five steps, in order
- Where the rest of turn taking lives - interruption, and the turn model
- Troubleshooting - reading
user_turn
agent.yaml
local/silero is what Pipecat runs. LiveKit needs its own turn model, so a
package that ships to both overrides it in targets.yaml:
targets.yaml
pace is the ceiling: the longest the agent will keep waiting before it
answers regardless. Reach for this first.
endpointing_delay is the floor: how long silence has to last before the
agent believes you finished. Nothing downstream can start earlier.
pace
snappy | balanced | patient
default:"balanced"
Legal on a
turn binding. The ceiling: the longest the agent will keep
waiting before it answers regardless. It sets the floor too when
endpointing_delay is absent. No per-target override.balanced when you leave it out means a package
that says nothing still gets the faster behaviour.
patient is the escape hatch. Selecting it changes nothing for a package that
already has today’s behaviour, so it is always safe to fall back to.
The two targets are not equally capable
unmute compile writes the resolved floor and ceiling for each target to
build/<target>/compile-report.json and the emitted README.md, so you never
have to look either up. Two differences are worth knowing before you read them:
- The Pipecat floor never moves. It is the same 0.2s at every pace, deliberately. See the cliff below.
- Only LiveKit adapts. It shortens its wait based on the pauses a caller actually leaves, between a lower bound and the ceiling. Pipecat has no equivalent, so its wait is the same fixed window every turn.
pace takes no per-target override. One word is meant to work on both
targets, and a value that differed per target would be a duration in disguise,
which is what endpointing_delay is for. Writing a pace inside a
targets.yaml override is refused, naming what to do instead.
endpointing_delay
positive duration
Legal on a
turn binding. The floor: the window of silence before the caller
counts as finished. Optional, and leaving it out lets the pace set the floor
too. LiveKit refuses anything under 250ms. Takes a per-target override.unmute compile rejects it rather than letting the worker raise on its first
call. But even a legal short value splits utterances: a pause between words
gets read as the end of the turn, and the agent answers before the sentence is
finished.
Unlike pace, this one does take a per-target override, because a silence
window is tied to one target’s mechanism. examples/salon-concierge authors no
floor on its base binding, so the pace owns it there, and 200ms on its Pipecat
target. The next section is why that one figure is authored rather than left to
the pace.
The Pipecat floor is a cliff, not a dial
Widening the Pipecat window can make turns slower. This is the one counter-intuitive thing on this page.pipecat-slng asks the bridge to finalise when voice activity detection reports
the caller stopped. A final transcript that has already arrived by then finds no
request outstanding, so the frame goes out unfinalized and Pipecat waits out a
flat safety-net timeout instead, on top of the window you set.
So the relationship between the window and the wait is not monotonic. Below the
transcript’s arrival time the turn ends promptly. Above it, the turn pays that
flat extra wait, and a transcript can arrive sooner than you would guess, so
“above it” starts sooner too.
That is why the Pipecat column in the pace table stays at 0.2s for every pace, and
why a patient Pipecat agent gets its patience from the ceiling instead. If you
raise this window on Pipecat, check the wait afterwards rather than assuming it
went up by what you added.
Setting interruption.minimum_words on a Pipecat target used to replace the
floor and ceiling with a plain timeout, which stopped the end-of-turn
classifier running and made turn taking worse instead of better. It no longer
does: the classifier survives, and it still carries the pace ceiling.
Let the transcriber decide
On Pipecat there is a third setting, and it changes what the other two mean. A Deepgram Flux or Cartesia Turns listener can end the turn itself: writeprovider: listen on the turn: binding. The pace ceiling then becomes the
transcriber’s own end-of-turn timeout (eot_timeout_ms or
turn_end_timeout_ms, in milliseconds), there is no floor because no local
silence window ends a turn, and endpointing_delay is refused.
Add eager: true and the reply is generated while the transcriber is still
confirming the caller stopped. The framework holds it and drops it if the
caller goes on or the confirmed words differ. It costs one model request per
prediction, including the withdrawn ones, so it is off unless you ask.
Turn detection has the shape, the two vendors, and
every refusal.
Reading what you set
Each target’s emittedbuild/<target>/README.md names the resolved pace, the
floor, and the ceiling, so you never have to infer them from the generated
code. It also says whether the floor came from your own endpointing_delay or
from the pace, and, on Pipecat, which of the two identically named stop_secs
fields in bot.py is the floor and which is the ceiling.
How to tune it
The order below matters. Each step tells you whether the next one is worth doing, and the first two need no audio at all.1. Read what you already have
2. Decide from the caller, not from the clock
Pick the pace from what your callers actually say, before measuring anything:- Do they answer in a few words: confirmations, a menu choice, a yes? Start at
snappy. - Do they read things out: phone numbers, dates, postcodes, an email address,
a name being spelled? Start at
balanced, and be ready to go topatient. - Do they think aloud, with pauses inside a sentence?
patient.
3. Listen before you measure
- Finish sentences cleanly and stop. This is the case a shorter ceiling improves, so it is where you will feel the change.
- Read a phone number aloud in groups, with a pause between each group. Then pause mid-sentence and carry on. If the agent answers your first group, or answers half your sentence, the pace is too fast for these callers. Go up one.
4. Then read the numbers
unmute dev reports the wait as its own number, user_turn, separately from
everything else in the turn. That separation is the point: it tells you whether
turn taking is your problem before you change anything.
Compare each turn’s user_turn against the floor and ceiling from step 1.
Troubleshooting has what each comparison means and what to
do about it.
5. When turn taking is not the answer
This is the common outcome once the ceiling is set sensibly. Look at how many model round trips the slow turns took. A turn that calls a tool pays time-to-first-token twice: once to decide the tool, once to answer with its result. Tool turns cost roughly double, and two tools in one turn cost more again. That is usually a bigger number than anything on this page, and no pace will touch it. The fixes are structural: collapse two tools into one, ask for one piece of information instead of three, or answer from context rather than looking something up. See Optimizing your agent for that side.Where the rest of turn taking lives
Everything on this page is on theturn binding under models. Two related
things are not, and it is worth knowing why:
conversation.interruptiondecides who holds the floor while the agent is speaking: whether a caller can barge in, how many words it takes, and which stretches of the call are protected. It is a conversation policy rather than a property of the turn detector, so it sits underconversation. Its fields are in the agent.yaml reference.- The turn model itself is per-target vendor selection, so a LiveKit package
names
turn-detector-miniin itstargets.yamloverride. See Turn detection for what actually runs on each target.
alpha and unlikely_threshold,
Pipecat’s pre_speech_ms and VAD confidence, are not reachable from a
package today. pace and endpointing_delay are the whole surface.
Troubleshooting
user_turn sits at your floor, turn after turn
The floor is the only thing being waited on. Turn taking is working.
Fix: nothing. Look at the model instead.
user_turn sits at your ceiling
The turn ran out of patience rather than deciding. These are the turns a shorter
ceiling saves.
Fix: go down one pace, then repeat
step 3.
user_turn is above your ceiling
Something else is holding the turn open. Check the transcription number on the
same turn: the ceiling cannot fire before the transcript arrives.
Fix: look at the transcriber, not the pace.
user_turn looks right and replies are still slow
Turn taking is not your problem.
Fix: see
When turn taking is not the answer.
Widening the Pipecat window made turns slower
The Pipecat floor is a cliff. Above the transcript’s arrival time the turn pays a flat extra wait on top of the window you set. Fix: come back down, and check the wait afterwards rather than assuming it moved by what you added. See The Pipecat floor is a cliff, not a dial.unmute compile warns that a turn params: block reaches nothing
A params: block on a turn binding is not the escape hatch. It is accepted
for shape but reaches neither framework, so unmute compile warns rather than
letting you believe it worked:
agent_id and fallback on a turn binding.
Fix: remove the block. pace and endpointing_delay are the whole surface.
Next
Execution Layer
Caching and routing for speech, on SLNG’s own layer.
Reading the latency numbers
Read
user_turn against the floor and ceiling you just set.