Voice framesbeta

The realtime frames and binary audio format a voice session uses, and the events it writes into the conversation log.

Voice reuses the chat WebSocket (wss://api.kletso.ai/v1/realtime). JSON frames control the session; audio travels as binary frames.

Client → server

framefields
voice.startmode: "vad" | "ptt", sampleRate: 24000, clientId
voice.stopreason: "user" | "background"
voice.commitpush-to-talk release: answer what was said
voice.playeditemId (or "current"), ms played so far; sent every 500 ms while the assistant speaks and on interruption
voice.texttext typed during the session; the voice model answers aloud

Binary frames

Byte 0 is the kind: 0x01 microphone audio (client → server), 0x02 assistant audio (server → client). The rest is PCM, 16-bit little-endian, mono, 24 kHz. Frames are at most 16 KB (20–60 ms is a good size). They are never stored, replayed or acknowledged; they bypass the seq log entirely. The SSE transport cannot carry them.

Events written to the conversation log

typedata
voice.startedvoiceSessionId, model, voice, limits { maxSeconds, idleSeconds }
voice.statestate: listening | thinking | speaking | idle
voice.transcriptrole: user, text, final
voice.interrupteditemId, audioEndMs
voice.endedreason: user | idle | limit | error | not_configured, usage { audioInSeconds, audioOutSeconds, costMicros }
avatar.moodmood, source: rule | agent | heuristic, ttlMs?

Spoken turns also produce the ordinary message.created, message.delta and message.completed events with modality: "voice", and tool calls produce the ordinary tool.* and ui.* events, so clients that know nothing about voice still show the conversation correctly.

Errors

error events with code not_configured (no OpenAI key on the agent), voice_unavailable (voice off, or not on a WebSocket), voice_rate_limited, voice_provider_error. A conflict error frame answers a second voice.start while one session is running.

Last updated 2026-09-28 · Report an issue with this page