Skip to main content
The student can talk to Lala on a live voice call. The call is a WebSocket session. The client streams microphone audio up, and the server streams Lala’s voice back. The server also sends live transcripts of both speakers. OpenAPI cannot describe a WebSocket, so this page is the contract.

Connect

The token goes in the token query parameter, not in a header. Before each call, the server retrieves coaching examples and puts them in the instructions of the model. To start a call without these examples, add rag=off to the URL. Use this parameter to compare two calls in a test. The default is on. The parameter has no effect on a reconnect.
The server closes the socket with one of these codes when the call cannot start:

Handshake

  1. Open the socket.
  2. Send the optional context message immediately. Do not wait for a prompt.
  3. Wait for the ready message.
  4. Start the microphone and send audio messages.
The server waits a maximum of 1 second for the context message. After that, the call starts without context. Audio that arrives before ready is dropped.

Messages from the client

Every message is a JSON text frame.
optional, first message only
Free text about why the student calls. The server puts it in the system prompt.
microphone audio
Base64 of raw 16-bit little-endian PCM at 16 kHz, mono. Send chunks of 20-100 ms.Stream continuously. The model detects the end of a turn from the silence in the audio it receives, so a stop in the stream holds the turn open. If the student mutes the microphone, send silence. Do not stop the stream.
optional, any time
The client heard the student stop speaking. The server tells the model to answer now. The model does not wait for its own silence timer.This message is optional. If the client does not send it, the model closes the turn after 1200 ms of silence. A wrong signal costs one early answer, not a lost reply.Send it after the audio chunk that ends the speech. If no new audio arrived since the last signal, the server ignores the message.If ready reports "hybridVad": false, do not send this message. Keep your detector off and let the model close the turn.
hang up
The student ended the call. The server closes the socket with code 1000.
optional, any time
Turns the diagnostics stream on or off for this call. See Call diagnostics. Send it again after a reconnect; the new socket starts with the stream off.

Messages from the server

call started
The voice session is up. Start the microphone now. callId identifies the call record.hybridVad says whether this server wants your own end-of-speech detector. If it is true, send speech_end when your detector hears the student stop. If it is false or absent, keep the detector off. The server can turn the feature off for every client without an app release.rag.enabled repeats the rag parameter of the URL. rag.chunks is the number of coaching examples that the server gave to the model. The number can be 0 when rag.enabled is true: no example was near enough, or the search was too slow.
Lala's voice
Base64 of raw 16-bit little-endian PCM at 24 kHz, mono. Note that the output rate (24 kHz) is not the input rate (16 kHz). Queue the chunks and play them in order.
barge-in
The student spoke over Lala. Flush the playback queue immediately. If you keep playing the queued audio, Lala talks over the student.
live captions
Incremental transcript of one speaker. speaker is user or lala. Chunks are deltas — concatenate them per speaker. Do not rely on final: true to close a segment: the upstream model rarely sets it. The reliable seams are conversational: Lala’s reply starting closes the student’s segment, and turn_complete or interrupted closes Lala’s.
turn boundary
Lala finished a full response turn: everything she wanted to say for the student’s last utterance has been streamed (playback may still be draining). Close her transcript segment.
call is over
The call ends on the server side. reason is model_ended (Lala said goodbye and hung up) or max_duration (the 30-minute limit). After model_ended, audio can continue for a short time — that is the goodbye. The socket closes after it.
non-fatal error
A server-side error report. The call continues when possible.

Call diagnostics

Turn-taking lives on the server and in the voice model, so a client that hears silence cannot tell why. The debug message opens a second stream of frames that says what the server sees. It is meant for the in-app call debugger, not for product features: the shape of these frames can change without notice.
sent once, on enable
Every diagnostic event the server recorded since the call started (the last 300). Enabling the stream after a problem still shows what led up to it.
one event
One diagnostic event. t is milliseconds since the call started, on the server clock. ev names the event; the other fields depend on it. Transcript events carry character counts, never text.
every second
A snapshot of the server side of the call: the inferred turn state, the server’s own energy VAD over the received microphone audio, how long ago the last microphone chunk arrived, the state of the upstream model socket, and the result of the search for coaching examples.
This is a rag event:
The events, in the order a normal exchange produces them: no_reply and empty_turn are also written to the server log for every call, with or without the stream enabled.

Reconnect after a network drop

The call survives a short network drop, for example a switch from WiFi to cellular. When the socket drops without an ended message, the server keeps the call alive for 25 seconds.
  1. Open a new socket to the same URL, with the callId from the ready message added:
  1. Wait for ready. On a resumed call it contains "resumed": true. Do not send context again — the prompt is set at call start.
  2. Continue to send audio messages.
Microphone audio that you send in the seconds around the drop is not lost. The server holds it and replays it to the model. If the server closes with code 4004, the grace period is over or a different server instance answered. Start a new call in that case. The server also sends a WebSocket protocol ping every 20 seconds. WebSocket libraries answer pings automatically — you do not have to write code for this. A client that does not answer two pings is disconnected and gets the 25-second reconnect window.

After the call

The server stores the full transcript with the call record. It also writes a summary into Lala’s conversation memory. As a result, the next chat message knows what was said on the call. The client does not have to send anything for this.

Limits

  • A call stops after 30 minutes.
  • A network drop longer than 25 seconds ends the call.
  • One short reconnect gap is possible mid-call while the server renews its model session. The socket to the client stays open through it, and microphone audio from the gap is replayed.