Smart Turn
A voice agent should respond when a person has finished a thought—not simply when the waveform becomes quiet.
Builder · Model adaptation, bilingual evaluation, and ONNX inference
93.9%
Hindi accuracy
93.7%
English accuracy
~38ms
per clip on CPU
Problem
Silence alone is an unreliable handoff signal for voice agents. People pause mid-thought, speak at different rates, and use language-specific phrasing that can make a fixed timeout either interrupt too early or respond too late.
Why it matters
Turn completion sits directly on the interaction loop. A small, bilingual model that runs quickly on CPU can improve responsiveness without adding another remote inference dependency.
Architecture
6 stages · select to inspect
Figures from the work


Technical challenges
Recognizing intent beyond silence
The classifier uses encoded speech context rather than a fixed pause threshold, allowing it to distinguish a completed utterance from a hesitation.
Holding quality across languages
Hindi and English are evaluated separately so aggregate accuracy cannot hide language-specific regressions.
Keeping the interaction loop fast
A compact Whisper-tiny backbone and single ONNX export keep inference local and predictable on CPU.
Tradeoffs
Whisper-tiny over a larger speech encoder
Turn detection needs useful representations at interaction latency; a larger backbone would increase the cost of every conversational handoff.
Eight-second context window
The window preserves enough recent speech structure for a completion decision while bounding inference work.
Single ONNX graph over a multi-stage runtime
One deployable graph reduces serving overhead and keeps preprocessing-to-decision latency easier to measure.
Experiments
- 01Evaluated turn-completion accuracy independently on Hindi and English clips.
- 02Measured CPU latency on the exported eight-second-context ONNX model.
- 03Validated the encoder, attention pooling, and classification head as one inference graph.
Results
Reached 93.9% accuracy on Hindi and 93.7% on English.
Achieved approximately 38ms inference per clip on CPU.
Exported the full turn-detection path as a single ONNX graph.
Lessons learned
- Conversational latency depends on deciding when to listen as much as how fast the response model runs.
- Per-language evaluation is necessary for a bilingual interaction model.
- A compact pretrained encoder can be adapted to a narrow product decision without carrying a full transcription pipeline.
Future work
- Evaluate noisier environments, accents, and longer conversational pauses.
- Measure false interruption and delayed-response costs in a live voice-agent loop.
- Calibrate confidence thresholds for different interaction styles.