When Is a Swedish Speaker Actually Done?
A voice AI can get every word right and still speak too soon. After “five millimeters,” a clinician may be finished or about to add the next measurement. The system has to know when to keep listening before it hears what comes next.
We built Dentio's Swedish end-of-turn model around that decision. In our internal Swedish evaluation, Dentio scored above every tested open model on MCC at 300 ms, using the comparators' published settings.
TLDR: End-of-turn detection is about knowing when to keep listening, not just recognizing a finished phrase. At 300 ms, Dentio's Swedish model reached 0.382 MCC, responded to 79% of YIELD moments, and made 20% fewer false response decisions during HOLD than hosted LiveKit Full in our evaluation.
The framework: learn a bounded correction
A silence timer measures how long a pause lasts. An endpoint model also uses the speech leading into it. Our framework, Anchor-Preserving Residual Fusion (APRF), gives Swedish conversational context a controlled influence on that endpoint estimate.
APRF keeps a fixed endpoint model as its anchor. A lightweight prediction head learns from frozen audio representations of Swedish speech. Its contextual estimate moves the anchor through a bounded residual correction.
The bound applies in logit space, before the score becomes a probability. Even a very confident contextual prediction can move the anchor only by a fixed maximum amount. We can train and inspect this correction separately, and disabling it returns the anchor's score exactly.
Figure 1 · Model architecture
How APRF adds Swedish context
The decision at 300 ms
We evaluate the decision our product makes: 300 ms after speech stops, is a response appropriate? Human reviewers mark whether to keep listening (HOLD), respond (YIELD), or allow either choice (OPTIONAL). A completed phrase can still be HOLD; the judgment depends on conversational context.
We summarize the decisions with Matthews correlation coefficient (MCC), which considers correct and incorrect decisions for both HOLD and YIELD. Always waiting or always responding scores zero. We also examine false response decisions and missed opportunities to respond.
In the same checkpoint, UltraVAD reached 0.289 MCC.
Figure 2 · Swedish conversations
Knowing when to keep listening
Swedish turn-taking
Our production use case
Published comparator rules. Dentio uses its frozen development threshold.
0 = no association · 1 = perfect agreement
Closed-source model
Human-reviewed Swedish pauses, with audio excluded from fitting; development can share recordings. Every model receives audio through 300 ms after speech stops. Processing time, network time and OPTIONAL judgments are excluded.
In the main evaluation, Dentio made 20% fewer false response decisions during HOLD than hosted LiveKit Full, while responding to 79% of YIELD opportunities at 300 ms. Full responded on more YIELD moments and scored higher on MCC. For our clinical workflows, we place greater weight on avoiding interruptions than on taking every opportunity to respond.
From a score to a conversational turn
The endpoint score feeds a turn controller. The controller can require sustained evidence, apply a minimum wait, and cancel a pending handover when speech resumes. This lets us change the interruption budget without retraining the model.
We apply the same separation to the audio pipeline: an audio chunk can end while the conversational turn stays open. Transcription can process that chunk while the turn controller keeps listening. This lets the system make progress through a hesitation without treating every chunk boundary as permission to respond.
The English reference
EOT-Bench measures how quickly a system can respond after a true turn ending without cutting the speaker off during earlier pauses. Its policy search varies the score threshold, minimum wait and timeout.
The English comparison below reports false cutoffs and average waiting time under that protocol. Here 300 ms is an average waiting budget after true endings. English provides a complementary check: a model specialized for Swedish should retain strong turn-taking behavior on a broader benchmark. Our Swedish specialization retains strong English performance, with fewer false cutoffs than the open models shown here at both the 300- and 600-ms waiting budgets.
Figure 3 · English EOT-Bench
The English comparison
Choose a waiting budget or false-cut threshold. Systems are ordered by the selected measure; lower is better.
False cutoff rate at a 300 ms average-wait budget↓ Lower is better
Closed-source model
We are building clinical software that can follow a conversation as it happens: recognize the words, hold context through a hesitation, and act at the right moment. If you want to work across speech models and the systems that put them to use, explore engineering opportunities at Dentio.
Our model development uses authorized, non-patient speech. Dentio patient information is not used to train our AI models.