Back to feed
Dev.to
Dev.to
8/2/2026
Why waiting longer makes voice AI worse

Why waiting longer makes voice AI worse

Short summary

Silence-based voice activity detection for turn-ending in voice AI creates a no-win tuning problem: short thresholds interrupt users mid-thought, long thresholds make the agent feel broken. The proposed alternative uses a lightweight endpointing model over partial transcripts that incorporates syntax, prosody, and semantics rather than relying solely on silence timing. Decoupling detection from commitment allows cheap cancellation when users resume, converting a hard classification into a soft one.

  • Silence-only VAD threshold tuning fails at both ends of the range
  • Better endpointing uses syntax, prosody, and semantics as signals
  • Decoupling detection from commitment enables cheap cancellation on false endpoints

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more