Voice framesbeta
The realtime frames and binary audio format a voice session uses, and the events it writes into the conversation log.
Voice reuses the chat WebSocket (wss://api.kletso.ai/v1/realtime). JSON frames control the session; audio travels as binary frames.
Client → server
| frame | fields |
|---|---|
voice.start | mode: "vad" | "ptt", sampleRate: 24000, clientId |
voice.stop | reason: "user" | "background" |
voice.commit | push-to-talk release: answer what was said |
voice.played | itemId (or "current"), ms played so far; sent every 500 ms while the assistant speaks and on interruption |
voice.text | text typed during the session; the voice model answers aloud |
Binary frames
Byte 0 is the kind: 0x01 microphone audio (client → server), 0x02 assistant audio (server → client). The rest is PCM, 16-bit little-endian, mono, 24 kHz. Frames are at most 16 KB (20–60 ms is a good size). They are never stored, replayed or acknowledged; they bypass the seq log entirely. The SSE transport cannot carry them.
Events written to the conversation log
| type | data |
|---|---|
voice.started | voiceSessionId, model, voice, limits { maxSeconds, idleSeconds } |
voice.state | state: listening | thinking | speaking | idle |
voice.transcript | role: user, text, final |
voice.interrupted | itemId, audioEndMs |
voice.ended | reason: user | idle | limit | error | not_configured, usage { audioInSeconds, audioOutSeconds, costMicros } |
avatar.mood | mood, source: rule | agent | heuristic, ttlMs? |
Spoken turns also produce the ordinary message.created, message.delta and message.completed events with modality: "voice", and tool calls produce the ordinary tool.* and ui.* events, so clients that know nothing about voice still show the conversation correctly.
Errors
error events with code not_configured (no OpenAI key on the agent), voice_unavailable (voice off, or not on a WebSocket), voice_rate_limited, voice_provider_error. A conflict error frame answers a second voice.start while one session is running.