Skip to content
Selected work
Audio intelligence · Voice agents · 2026

Smart Turn

A voice agent should respond when a person has finished a thought—not simply when the waveform becomes quiet.

Builder · Model adaptation, bilingual evaluation, and ONNX inference

93.9%

Hindi accuracy

93.7%

English accuracy

~38ms

per clip on CPU

01

Problem

Silence alone is an unreliable handoff signal for voice agents. People pause mid-thought, speak at different rates, and use language-specific phrasing that can make a fixed timeout either interrupt too early or respond too late.

02

Why it matters

Turn completion sits directly on the interaction loop. A small, bilingual model that runs quickly on CPU can improve responsiveness without adding another remote inference dependency.

03

Architecture

6 stages · select to inspect

The model reuses Whisper-tiny's speech representation, aggregates variable acoustic context with attention pooling, and emits a turn-completion decision through a lightweight classification head. The complete path is exported as one ONNX graph for CPU inference.
04

Figures from the work

Smart Turn: waveform and log-mel spectrogram of a completed turn beside one where the speaker is still talking, with the trailing pause marked
Smart Turn dataset analysis: language balance, real versus synthetic sources, speech duration, trailing pause by class, and signal-to-noise by source
05

Technical challenges

01

Recognizing intent beyond silence

The classifier uses encoded speech context rather than a fixed pause threshold, allowing it to distinguish a completed utterance from a hesitation.

02

Holding quality across languages

Hindi and English are evaluated separately so aggregate accuracy cannot hide language-specific regressions.

03

Keeping the interaction loop fast

A compact Whisper-tiny backbone and single ONNX export keep inference local and predictable on CPU.

06

Tradeoffs

Whisper-tiny over a larger speech encoder

Turn detection needs useful representations at interaction latency; a larger backbone would increase the cost of every conversational handoff.

Eight-second context window

The window preserves enough recent speech structure for a completion decision while bounding inference work.

Single ONNX graph over a multi-stage runtime

One deployable graph reduces serving overhead and keeps preprocessing-to-decision latency easier to measure.

07

Experiments

  1. 01Evaluated turn-completion accuracy independently on Hindi and English clips.
  2. 02Measured CPU latency on the exported eight-second-context ONNX model.
  3. 03Validated the encoder, attention pooling, and classification head as one inference graph.
08

Results

Reached 93.9% accuracy on Hindi and 93.7% on English.

Achieved approximately 38ms inference per clip on CPU.

Exported the full turn-detection path as a single ONNX graph.

09

Lessons learned

  • Conversational latency depends on deciding when to listen as much as how fast the response model runs.
  • Per-language evaluation is necessary for a bilingual interaction model.
  • A compact pretrained encoder can be adapted to a narrow product decision without carrying a full transcription pipeline.
10

Future work

  • Evaluate noisier environments, accents, and longer conversational pauses.
  • Measure false interruption and delayed-response costs in a live voice-agent loop.
  • Calibrate confidence thresholds for different interaction styles.