Skip to main content

Voice & turn-taking

Interrupt Finn.Finn yields.

Natural turn-taking — when the caller speaks over Finn, Finn stops mid-word. No talking over each other. No awkward pauses.

Book a demo

30-minute free trial. No card. No procurement.

<80msYield latency
NaturalMid-word handoff
No clippingResumes cleanly
Per-turnTunable thresholds

How it works

Listen evenwhile speaking.

Drag, wire, deploy. Each step is reversible — every save creates a versioned snapshot.

01 / 03

Bidirectional audio stream

Finn's mic stays open during synthesis. VAD detects caller voice the instant they start.

Deployment analyticsSub-second voice loopHealthy

p50

384ms

p95

412ms

p99

538ms

ASRDeepgram Nova 2 streaming
84ms
LLMGPT-4o function-calling pass
192ms
TTSElevenLabs streaming synthesis
96ms
Last 24h · 12,840 calls384ms avg
02 / 03

Yield mid-word

Under 80ms from detected voice to TTS stop. No mid-sentence collisions.

Deployment analyticsSub-second voice loopHealthy

p50

384ms

p95

412ms

p99

538ms

ASRDeepgram Nova 2 streaming
84ms
LLMGPT-4o function-calling pass
192ms
TTSElevenLabs streaming synthesis
96ms
Last 24h · 12,840 calls384ms avg
03 / 03

Resume context-aware

Acknowledges interruption, processes what the caller said, picks up the right thread. No 'as I was saying.'

OrchestrationMulti-agent flowLive
T
01

Triage

Inbound router

DONE
B
02

Booking

Scheduling specialist

ACTIVE
C
03

Confirm

SMS follow-up

QUEUED

Shared context · call IN-44219

caller.name      = "Aarav Singh"
caller.intent    = "follow-up booking"
captured.slot    = "Tue 14:00"
captured.doctor  = "DR. Patel"
captured.insurer = "BCBS"

Reads like real conversation

Caller leads.Finn follows.

No bolt-ons. No second platform. Every primitive ships with the editor.

Sub-80ms yield

Faster than human reaction time. Finn drops the floor before the caller registers they're talking over.

Clean resume

Tracks what was said + what was cut off. Acknowledges + adapts. No robotic 'one moment please.'

Tunable barge-in threshold

Per-workflow VAD sensitivity. Quiet caller in a quiet room vs busy contact center — both work.

False-positive resistant

Trained to distinguish caller voice from ambient noise, music, cross-talk. Finn doesn't yield to a sneeze.

Shipped in production

Real teams, real outcomes.

TOFA launched a national campaign with Finn — 10K+ calls daily and seamless human escalations.
Alice Smith

Alice Smith

Senior Engineer, Gofts · TOFA

Trust + ecosystem

Plugs in. Audits clean.

Wires into your existing stack via 100+ connectors. Compliant from day one across the workloads that demand it.

Twilio logoTwilioPlivo logoPlivoElevenLabsOpenAIMicrosoft Teams logoMicrosoft Teams

Don't see yours? Finn plugs into any system via REST + webhooks.

AuditedSOC 2 Type IIHIPAA BAAGDPRISO 27001DPDPNIST CSF

FAQ

Questions and answers.

What's the technical latency for barge-in?+

Sub-80ms from detected voice onset to TTS halt. Faster than the caller can perceive their own interruption.

How does Finn handle the interruption content?+

Captures what the caller said, processes it against the conversation context, and responds appropriately — apology, clarification, branch to new topic, whatever fits.

Can I tune sensitivity per workflow?+

Yes. Quiet inbound support calls use a low threshold. Noisy outbound to mobile users uses a higher threshold to ignore ambient noise.

What about cross-talk on bad connections?+

Echo cancellation + caller-voice classification filter cross-talk + line noise. Finn doesn't yield to its own voice bouncing back.

Does this work across all telephony providers?+

Yes — Twilio, Plivo, BYOC SIP trunk. Requires bidirectional streaming audio, which all major providers support.

Conversation, not monologue

Yields like a human.Not a kiosk.

Real-time barge-in detection — Finn drops the floor the instant the caller speaks. No more 'press 1 to interrupt.'

Book a demo30-min trial · No card