CAMLIN SPEECH

Streaming STT, Australian voices, and frame-accurate lip-sync

Camlin Speech streams words onto the screen as you speak and ends the turn on a natural pause. Camlin Voice answers with Sydney-recorded Australian voices and avatar lip-sync in the same response. Our engine by default; Google, Deepgram, Azure, and others when a tender needs them.
5
Speech providers
Camlin default; swap per tender
8
Recognition modes
Spoken to structured
16
Business field types
Dates, numbers, IDs
Try the live demo

Opens the speech portal — mic on, your words land as you speak.

Camlin Speech

Recognition, prompting, and validation

LIVE EXAMPLE
Recognition result

Recognition + prompting
Capture, validate, and respond in one service
Structured output
Business fields and confidence checks
THE SPEECH STACK

Own the speech layer, keep provider choice

Recognition, TTS, prompts, visemes, and structured capture live in one layer so Contact, Voice, Avatar, and every channel can share the same speech logic.

Streaming STT, ended on a natural pause

Words commit on screen as the caller speaks. The recogniser closes the turn on a natural pause — not a fixed silence timer — and the endpointing is authorable per Interaction, so a reference-number capture can wait longer than a yes/no.

Word-by-word live captions
Authorable natural-pause endpointing
Sydney or private deployment

Tenant vocabularies and field capture

Register the names, brands, reference formats, and service terms your callers will say. The speech layer biases recognition and routes output into business fields instead of raw transcript text.

Per-tenant lexicons
Names, IDs, dates, and references
Structured output for Interactions

Camlin Voice: TTS with lip-sync built in

Every synthesis response carries the audio plus a 52-blendshape ARKit track rendered at 60 fps — the same envelope that drives the 3D avatar’s mouth on a live call. No Azure dependency in the lip-sync path.

52 ARKit blendshapes at 60 fps
Audio + lip-sync in one response
AU voices recorded in Sydney

Provider choice when you need it

The Camlin engine is the default. Google, Deepgram, Azure, and AWS stay available per journey for tenders, specific languages, or fallback — authored in Architect without splitting the service.

Google, Deepgram, Azure, AWS support
Provider per journey
8 recognition modes
SPEECH PIPELINE

How speech flows through the platform

Follow spoken or typed input through recognition, validation, and response. Each step connects to a different speech capability.

Caller speaks or types

Audio or text input arrives from the channel — phone, web chat, avatar, or SMS.

The channel determines which speech provider and mode to use based on the Interaction design.

CAMLIN VOICE

Australian voices, frame-accurate lip-sync

Camlin Voice generates the audio and the blendshape stream that drives the avatar’s mouth in lock-step — one engine end-to-end, no third-party lip-sync dependency.
Recorded in Sydney

First-party Australian voices, recorded in our own studio. The roster keeps growing; every voice ships with the same lip-sync envelope.

Lip-sync in the same response

52 ARKit blendshapes at 60 fps ride along with the audio — the exact track the Nua avatar consumes on a live call.

Yours to build on

The waveform and viseme track come back as JSON, so the same call that voices your IVR can drive your own avatar pipeline.

HEAR IT LIVE

Hear the difference

Side-by-side clips where Camlin Speech nails Australian phrases the cloud STTs drop. Then try the live engine on your own voice.

More live demos on the speech portal: speech.camlinconnect.com.au. Prefer to stay on this site? See the embedded caption preview.

What’s next for Camlin Speech

Streaming captions on this site’s companion mic · liveMore Sydney-recorded AU voices · in studioPlain-WebSocket streaming API · planned