Voice and avatarbeta

Turn on voice for an agent with your own OpenAI key, pick a voice and limits, and shape how the avatar looks and reacts.

Both live on the agent, next to Persona and Model, and are versioned and published like everything else on the agent.

Voice tab

  • Voice on/off. Off hides the microphone in every app using this agent.
  • Model. gpt-realtime-2.1 (best quality), gpt-realtime-2.1-mini (about a third of the price) or gpt-realtime-2. The price line under the select is an estimate per spoken minute; it is paid to OpenAI on your key.
  • Voice. Ten OpenAI voices; marin and cedar sound best.
  • Your OpenAI key. A project secret. Until one is selected (or when the selected secret is deleted) voice sessions end immediately with not_configured, and the Overview page shows an “Add your API key” card with an OpenAI row that fixes every voice-enabled agent at once.
  • Speaking style. Extra instructions for the voice model only (“warm and concise, one or two sentences, offer to show things on screen”).
  • Turn detection. Natural (semantic), Fast (server VAD) or Push-to-talk (the user holds the mic).
  • Transcribe what the user says. On by default so the conversation history is complete; adds a small per-minute cost.
  • Tools allowed in voice. All of the agent’s tools by default; untick the ones that should stay text-only. Built-in tools (render UI, notify, app command, handoff) are always available.
  • Limits. Maximum minutes per session (1–55) and seconds of silence before the session ends.

Avatar tab

  • Style. Kletso mascot, your image, or none (the plain chat bubble icon).
  • Your image, per mood. With the image style you set a default picture (upload or https URL) and, on the Allowed moods strip, an optional picture per mood: Upload under a face adds one, × returns that mood to the default. PNG, JPG, GIF or WebP up to 2 MB; animated GIF/WebP play as-is. Uploads live in Kletso storage under a content-hashed https://api.kletso.ai/assets/… URL, so replacing a picture takes effect on the next publish with no caching surprises. The renderer animates the pictures itself: a crossfade when the mood changes, a slow breathing loop while idle, and a talk bounce from the voice level; the dashboard preview (“Play a conversation”) and the Flutter KletsoAvatar behave the same.
  • Colours. Inherit your widget branding, or set body, eyes, mouth and tongue.
  • Default mood and the set of allowed moods; the preview strip shows every face.
  • Rules. Event → mood → how long it holds. Defaults: typing or a tool running → thinking, a failed tool or an error → sorry, a finished answer → happy for two seconds, listening/speaking during voice, a wink in the launcher when a message arrives unread.
  • Let the agent pick moods. Adds a set_mood tool the model may call when the emotion is clear.
  • Preview. The avatar at full size with a “play a conversation” button: listening → thinking → speaking → happy.

Where it shows up

Conversations list spoken messages with a microphone mark and show the voice events (started, interrupted, ended with seconds in and out) in the timeline. Analytics gets a Voice card with sessions and minutes once the first voice session has run. Settings → Widget branding can set the launcher icon to the mascot.

Last updated 2026-09-28 · Report an issue with this page