Voice and the avatarbeta
How a Kletso agent talks and listens, what the animated avatar is, and what runs on whose key.
An agent can talk. The end user taps the microphone in the chat, speaks, and the assistant answers out loud while the Kletso avatar listens, thinks and lip-syncs. Everything the assistant can do in text it can do by voice: call your tools, render cards and forms on screen, send notifications, hand off to a person.
How a voice session works
- The app opens a voice session on the current conversation over the same WebSocket the chat already uses. Audio travels as small binary frames (PCM, 24 kHz).
- The Kletso runtime relays the audio to OpenAI’s realtime speech model (
gpt-realtime-2.1by default) on your own OpenAI key, stored as a project secret. Kletso never sees your key in the clear outside the conversation’s isolated worker, and never pays for or proxies model usage on its own account. - The model’s speech comes back as audio; its words come back as a transcript. The transcript is written into the conversation exactly like typed messages (marked
modality: voice), so the dashboard inspector, analytics and history all show what was said. - When the model decides to call a tool, the runtime runs it with the same executor as text chat: your HTTP tools, confirmations,
render_uisurfaces that appear under the avatar, notifications and app commands. - The user can interrupt at any moment; the assistant stops, the half-spoken sentence is cut in the transcript too, and listening resumes.
A session ends when the user taps End, after the configured silence, at the configured maximum length, or when the app goes to the background.
What it costs
Voice runs on your OpenAI account. The realtime model is billed per audio token; in practice a conversation costs about $0.06–0.11 per minute on gpt-realtime-2.1 and $0.02–0.05 on the mini model, plus about $0.006–0.017 per minute if user transcription is on (it is on by default so the chat transcript is complete). The dashboard shows voice minutes and the estimated cost per day next to the token counters.
The avatar
The avatar is the Kletso mascot’s face: a round, friendly character with two glossy eyes and an expressive mouth. It is not an image. The SDK draws it live, so it can breathe, blink, look around, blend between moods and move its mouth with the voice. The dashboard renders the same face from the same definition, so what you configure is what the app shows.
Moods: neutral, happy, laughing, surprised, thinking, wink, listening, speaking, sorry, confused, sleepy. They come from three sources, in this order of precedence:
- the agent itself, through the optional
set_moodtool (off by default: every tool call costs a round trip); - the runtime, which guesses a mood from what the assistant says during voice turns (“Sorry, …” → sorry, “Great!” → happy);
- client-side rules you configure per event: a tool running → thinking, an error → sorry, a finished answer → happy for two seconds, an unread message → a wink in the launcher.
You decide which moods are allowed, the default mood, the colours (inherit your widget branding or set your own), or replace the mascot with your own picture (it gets a speaking ring instead of a mouth) or with no face at all.
Where it is configured
Per agent, in the dashboard: the Voice tab (on/off, model, voice, your OpenAI key, speaking style, turn detection, which tools voice may call, limits) and the Avatar tab (style, colours, moods, rules, live preview). See Dashboard → Voice and avatar and Flutter → Voice.
Privacy
Kletso does not store audio. Transcripts are stored with the conversation like any message. Audio is processed by OpenAI under your account’s data policy. Tell your users that voice conversations are transcribed.