Skip to main content
    Tillbaka

    When Is a Swedish Speaker Actually Done?

    30 SEPTEMBER 2026•3 min läsning•By Fabian Farestam

    A voice AI can get every word right and still speak too soon. After “five millimeters,” a clinician may be finished or about to add the next measurement. The system has to know when to keep listening before it hears what comes next.

    We built Dentio's Swedish end-of-turn model around that decision. In our internal Swedish evaluation, Dentio scored above every tested open model on MCC at 300 ms, using the comparators' published settings.

    TLDR: End-of-turn detection is about knowing when to keep listening, not just recognizing a finished phrase. At 300 ms, Dentio's Swedish model reached 0.382 MCC, responded to 79% of YIELD moments, and made 20% fewer false response decisions during HOLD than hosted LiveKit Full in our evaluation.

    The framework: learn a bounded correction

    A silence timer measures how long a pause lasts. An endpoint model also uses the speech leading into it. Our framework, Anchor-Preserving Residual Fusion (APRF), gives Swedish conversational context a controlled influence on that endpoint estimate.

    APRF keeps a fixed endpoint model as its anchor. A lightweight prediction head learns from frozen audio representations of Swedish speech. Its contextual estimate moves the anchor through a bounded residual correction.

    The bound applies in logit space, before the score becomes a probability. Even a very confident contextual prediction can move the anchor only by a fixed maximum amount. We can train and inspect this correction separately, and disabling it returns the anchor's score exactly.

    Figure 1 · Model architecture

    How APRF adds Swedish context

    B limits the correction in logit space. Zero correction preserves the original anchor score exactly.

    The decision at 300 ms

    We evaluate the decision our product makes: 300 ms after speech stops, is a response appropriate? Human reviewers mark whether to keep listening (HOLD), respond (YIELD), or allow either choice (OPTIONAL). A completed phrase can still be HOLD; the judgment depends on conversational context.

    We summarize the decisions with Matthews correlation coefficient (MCC), which considers correct and incorrect decisions for both HOLD and YIELD. Always waiting or always responding scores zero. We also examine false response decisions and missed opportunities to respond.

    In the same checkpoint, UltraVAD reached 0.289 MCC.

    Figure 2 · Swedish conversations

    Knowing when to keep listening

    Swedish turn-taking

    Our production use case

    MCC · 300 ms
    Dentio (closed-source model)0.376
    LiveKit Full (closed-source model)0.264

    Main Swedish evaluation

    MCC · 300 ms ↑ Higher is better

    Published comparator rules. Dentio uses its frozen development threshold.

    LiveKit Full (closed-source model)0.465Correctly waits on 50.0% of HOLD decisions; responds on 94.7% of YIELD decisions.
    Dentio Swedish EOT (closed-source model)0.382Correctly waits on 60.0% of HOLD decisions; responds on 78.9% of YIELD decisions.
    TurnSense0.321Correctly waits on 100.0% of HOLD decisions; responds on 15.8% of YIELD decisions.
    UltraVAD0.289Correctly waits on 50.0% of HOLD decisions; responds on 78.9% of YIELD decisions.
    Namo + KB Whisper Tiny0.161Correctly waits on 23.3% of HOLD decisions; responds on 89.5% of YIELD decisions.
    LiveKit Mini0.084Correctly waits on 10.0% of HOLD decisions; responds on 94.7% of YIELD decisions.
    SmartTurn 3.20.064Correctly waits on 26.7% of HOLD decisions; responds on 78.9% of YIELD decisions.
    DualTurn0.007Correctly waits on 53.3% of HOLD decisions; responds on 47.4% of YIELD decisions.

    0 = no association · 1 = perfect agreement

    Closed-source model

    Human-reviewed Swedish pauses, with audio excluded from fitting; development can share recordings. Every model receives audio through 300 ms after speech stops. Processing time, network time and OPTIONAL judgments are excluded.

    In the main evaluation, Dentio made 20% fewer false response decisions during HOLD than hosted LiveKit Full, while responding to 79% of YIELD opportunities at 300 ms. Full responded on more YIELD moments and scored higher on MCC. For our clinical workflows, we place greater weight on avoiding interruptions than on taking every opportunity to respond.

    From a score to a conversational turn

    The endpoint score feeds a turn controller. The controller can require sustained evidence, apply a minimum wait, and cancel a pending handover when speech resumes. This lets us change the interruption budget without retraining the model.

    We apply the same separation to the audio pipeline: an audio chunk can end while the conversational turn stays open. Transcription can process that chunk while the turn controller keeps listening. This lets the system make progress through a hesitation without treating every chunk boundary as permission to respond.

    The English reference

    EOT-Bench measures how quickly a system can respond after a true turn ending without cutting the speaker off during earlier pauses. Its policy search varies the score threshold, minimum wait and timeout.

    The English comparison below reports false cutoffs and average waiting time under that protocol. Here 300 ms is an average waiting budget after true endings. English provides a complementary check: a model specialized for Swedish should retain strong turn-taking behavior on a broader benchmark. Our Swedish specialization retains strong English performance, with fewer false cutoffs than the open models shown here at both the 300- and 600-ms waiting budgets.

    Figure 3 · English EOT-Bench

    The English comparison

    Choose a waiting budget or false-cut threshold. Systems are ordered by the selected measure; lower is better.

    False cutoff rate at a 300 ms average-wait budget↓ Lower is better

    LiveKit Full (closed-source model)9.929%
    Deepgram Flux (closed-source model)12.9%
    Dentio Swedish EOT (closed-source model)22.837%
    ultraVAD27.7%
    LiveKit Mini27.801%
    SmartTurn 3.235.2%
    AssemblyAI (closed-source model)49.4%
    Silence-only VAD baseline55.6%

    Closed-source model

    Our English measurements alongside published EOT-Bench results. The 300 / 600-ms settings are average waiting budgets after true turn endings. This is separate from the Swedish decision test.

    We are building clinical software that can follow a conversation as it happens: recognize the words, hold context through a hesitation, and act at the right moment. If you want to work across speech models and the systems that put them to use, explore engineering opportunities at Dentio.


    Our model development uses authorized, non-patient speech. Dentio patient information is not used to train our AI models.