Dev.to
8/2/2026

Why waiting longer makes voice AI worse
Short summary
Silence-based voice activity detection for turn-ending in voice AI creates a no-win tuning problem: short thresholds interrupt users mid-thought, long thresholds make the agent feel broken. The proposed alternative uses a lightweight endpointing model over partial transcripts that incorporates syntax, prosody, and semantics rather than relying solely on silence timing. Decoupling detection from commitment allows cheap cancellation when users resume, converting a hard classification into a soft one.
- •Silence-only VAD threshold tuning fails at both ends of the range
- •Better endpointing uses syntax, prosody, and semantics as signals
- •Decoupling detection from commitment enables cheap cancellation on false endpoints
Generated with AI, which can make mistakes.
Is this a good recommendation for you?



