Connect
token query parameter, not in a header.
Before each call, the server retrieves coaching examples and puts them in the instructions of the
model. To start a call without these examples, add rag=off to the URL. Use this parameter to
compare two calls in a test. The default is on. The parameter has no effect on a reconnect.
Handshake
- Open the socket.
- Send the optional
contextmessage immediately. Do not wait for a prompt. - Wait for the
readymessage. - Start the microphone and send
audiomessages.
context message. After that, the call starts
without context. Audio that arrives before ready is dropped.
Messages from the client
Every message is a JSON text frame.optional, first message only
Free text about why the student calls. The server puts it in the system prompt.
microphone audio
Base64 of raw 16-bit little-endian PCM at 16 kHz, mono. Send chunks of 20-100 ms.Stream continuously. The model detects the end of a turn from the silence in the audio it
receives, so a stop in the stream holds the turn open. If the student mutes the microphone,
send silence. Do not stop the stream.
optional, any time
The client heard the student stop speaking. The server tells the model to answer now. The model
does not wait for its own silence timer.This message is optional. If the client does not send it, the model closes the turn after
1200 ms of silence. A wrong signal costs one early answer, not a lost reply.Send it after the audio chunk that ends the speech. If no new audio arrived since the last
signal, the server ignores the message.If
ready reports "hybridVad": false, do not send this message. Keep your detector off and
let the model close the turn.hang up
The student ended the call. The server closes the socket with code
1000.optional, any time
Turns the diagnostics stream on or off for this call. See Call diagnostics.
Send it again after a reconnect; the new socket starts with the stream off.
Messages from the server
call started
The voice session is up. Start the microphone now.
callId identifies the call record.hybridVad says whether this server wants your own end-of-speech detector. If it is true,
send speech_end when your detector hears the student stop. If it is false or
absent, keep the detector off. The server can turn the feature off for every client without an
app release.rag.enabled repeats the rag parameter of the URL. rag.chunks is the number of coaching
examples that the server gave to the model. The number can be 0 when rag.enabled is true:
no example was near enough, or the search was too slow.Lala's voice
Base64 of raw 16-bit little-endian PCM at 24 kHz, mono. Note that the output rate (24 kHz) is
not the input rate (16 kHz). Queue the chunks and play them in order.
barge-in
The student spoke over Lala. Flush the playback queue immediately. If you keep playing the
queued audio, Lala talks over the student.
live captions
Incremental transcript of one speaker.
speaker is user or lala. Chunks are deltas —
concatenate them per speaker. Do not rely on final: true to close a segment: the upstream
model rarely sets it. The reliable seams are conversational: Lala’s reply starting closes the
student’s segment, and turn_complete or interrupted closes Lala’s.turn boundary
Lala finished a full response turn: everything she wanted to say for the student’s last
utterance has been streamed (playback may still be draining). Close her transcript segment.
call is over
The call ends on the server side.
reason is model_ended (Lala said goodbye and hung up) or
max_duration (the 30-minute limit). After model_ended, audio can continue for a short time —
that is the goodbye. The socket closes after it.non-fatal error
A server-side error report. The call continues when possible.
Call diagnostics
Turn-taking lives on the server and in the voice model, so a client that hears silence cannot tell why. Thedebug message opens a second stream of frames that says what the server sees.
It is meant for the in-app call debugger, not for product features: the shape of these frames can
change without notice.
sent once, on enable
Every diagnostic event the server recorded since the call started (the last 300). Enabling the
stream after a problem still shows what led up to it.
one event
One diagnostic event.
t is milliseconds since the call started, on the server clock. ev
names the event; the other fields depend on it. Transcript events carry character counts,
never text.every second
A snapshot of the server side of the call: the inferred turn state, the server’s own energy
VAD over the received microphone audio, how long ago the last microphone chunk arrived, the
state of the upstream model socket, and the result of the search for coaching examples.
rag event:
no_reply and empty_turn are also written to the server log for every call, with or without the
stream enabled.
Reconnect after a network drop
The call survives a short network drop, for example a switch from WiFi to cellular. When the socket drops without anended message, the server keeps the call alive for 25 seconds.
- Open a new socket to the same URL, with the
callIdfrom thereadymessage added:
- Wait for
ready. On a resumed call it contains"resumed": true. Do not sendcontextagain — the prompt is set at call start. - Continue to send
audiomessages.
4004, the grace period is over or a different server instance
answered. Start a new call in that case.
The server also sends a WebSocket protocol ping every 20 seconds. WebSocket libraries answer
pings automatically — you do not have to write code for this. A client that does not answer two
pings is disconnected and gets the 25-second reconnect window.
After the call
The server stores the full transcript with the call record. It also writes a summary into Lala’s conversation memory. As a result, the next chat message knows what was said on the call. The client does not have to send anything for this.Limits
- A call stops after 30 minutes.
- A network drop longer than 25 seconds ends the call.
- One short reconnect gap is possible mid-call while the server renews its model session. The socket to the client stays open through it, and microphone audio from the gap is replayed.