> ## Documentation Index
> Fetch the complete documentation index at: https://docs.lala.ist/llms.txt
> Use this file to discover all available pages before exploring further.

# Voice calls

> The WebSocket contract of /api/call/ws: audio in, audio out, live transcripts.

The student can talk to Lala on a live voice call. The call is a WebSocket session. The client
streams microphone audio up, and the server streams Lala's voice back. The server also sends live
transcripts of both speakers. OpenAPI cannot describe a WebSocket, so this page is the contract.

## Connect

```text theme={null}
wss://client-api.lala.ist/api/call/ws?token=<supabase_jwt>
```

The token goes in the `token` query parameter, not in a header.

Before each call, the server retrieves coaching examples and puts them in the instructions of the
model. To start a call without these examples, add `rag=off` to the URL. Use this parameter to
compare two calls in a test. The default is on. The parameter has no effect on a reconnect.

```text theme={null}
wss://client-api.lala.ist/api/call/ws?token=<supabase_jwt>&rag=off
```

The server closes the socket with
one of these codes when the call cannot start:

| Close code | Meaning |
| - | - |
| `4001` | The token is missing or invalid. |
| `4002` | The server could not load the student. |
| `4003` | The server could not connect to the voice model. |
| `4004` | The `callId` in a reconnect attempt is unknown or expired. Start a new call. |

## Handshake

1. Open the socket.
2. Send the optional `context` message immediately. Do not wait for a prompt.
3. Wait for the `ready` message.
4. Start the microphone and send `audio` messages.

The server waits a maximum of 1 second for the `context` message. After that, the call starts
without context. Audio that arrives before `ready` is dropped.

## Messages from the client

Every message is a JSON text frame.

<ResponseField name="context" type="optional, first message only">
  Free text about why the student calls. The server puts it in the system prompt.

  ```json theme={null}
  { "type": "context", "context": "Öğrenci deneme sonucu ekranından aradı." }
  ```
</ResponseField>

<ResponseField name="audio" type="microphone audio">
  Base64 of raw 16-bit little-endian PCM at 16 kHz, mono. Send chunks of 20-100 ms.

  Stream continuously. The model detects the end of a turn from the silence in the audio it
  receives, so a stop in the stream holds the turn open. If the student mutes the microphone,
  send silence. Do not stop the stream.

  ```json theme={null}
  { "type": "audio", "data": "<base64>", "mimeType": "audio/pcm;rate=16000" }
  ```
</ResponseField>

<ResponseField name="speech_end" type="optional, any time">
  The client heard the student stop speaking. The server tells the model to answer now. The model
  does not wait for its own silence timer.

  This message is optional. If the client does not send it, the model closes the turn after
  1200 ms of silence. A wrong signal costs one early answer, not a lost reply.

  Send it after the audio chunk that ends the speech. If no new audio arrived since the last
  signal, the server ignores the message.

  If `ready` reports `"hybridVad": false`, do not send this message. Keep your detector off and
  let the model close the turn.

  ```json theme={null}
  { "type": "speech_end" }
  ```
</ResponseField>

<ResponseField name="end" type="hang up">
  The student ended the call. The server closes the socket with code `1000`.

  ```json theme={null}
  { "type": "end" }
  ```
</ResponseField>

<ResponseField name="debug" type="optional, any time">
  Turns the diagnostics stream on or off for this call. See [Call diagnostics](#call-diagnostics).
  Send it again after a reconnect; the new socket starts with the stream off.

  ```json theme={null}
  { "type": "debug", "enabled": true }
  ```
</ResponseField>

## Messages from the server

<ResponseField name="ready" type="call started">
  The voice session is up. Start the microphone now. `callId` identifies the call record.

  `hybridVad` says whether this server wants your own end-of-speech detector. If it is `true`,
  send [`speech_end`](#speech-end) when your detector hears the student stop. If it is `false` or
  absent, keep the detector off. The server can turn the feature off for every client without an
  app release.

  `rag.enabled` repeats the `rag` parameter of the URL. `rag.chunks` is the number of coaching
  examples that the server gave to the model. The number can be `0` when `rag.enabled` is `true`:
  no example was near enough, or the search was too slow.

  ```json theme={null}
  { "type": "ready", "callId": "9a8b...", "hybridVad": true, "rag": { "enabled": true, "chunks": 2 } }
  ```
</ResponseField>

<ResponseField name="audio" type="Lala's voice">
  Base64 of raw 16-bit little-endian PCM at 24 kHz, mono. Note that the output rate (24 kHz) is
  not the input rate (16 kHz). Queue the chunks and play them in order.

  ```json theme={null}
  { "type": "audio", "data": "<base64>", "mimeType": "audio/pcm;rate=24000" }
  ```
</ResponseField>

<ResponseField name="interrupted" type="barge-in">
  The student spoke over Lala. Flush the playback queue immediately. If you keep playing the
  queued audio, Lala talks over the student.

  ```json theme={null}
  { "type": "interrupted" }
  ```
</ResponseField>

<ResponseField name="transcript" type="live captions">
  Incremental transcript of one speaker. `speaker` is `user` or `lala`. Chunks are deltas —
  concatenate them per speaker. Do not rely on `final: true` to close a segment: the upstream
  model rarely sets it. The reliable seams are conversational: Lala's reply starting closes the
  student's segment, and `turn_complete` or `interrupted` closes Lala's.

  ```json theme={null}
  { "type": "transcript", "speaker": "lala", "text": "selam! ", "final": false }
  ```
</ResponseField>

<ResponseField name="turn_complete" type="turn boundary">
  Lala finished a full response turn: everything she wanted to say for the student's last
  utterance has been streamed (playback may still be draining). Close her transcript segment.

  ```json theme={null}
  { "type": "turn_complete" }
  ```
</ResponseField>

<ResponseField name="ended" type="call is over">
  The call ends on the server side. `reason` is `model_ended` (Lala said goodbye and hung up) or
  `max_duration` (the 30-minute limit). After `model_ended`, audio can continue for a short time —
  that is the goodbye. The socket closes after it.

  ```json theme={null}
  { "type": "ended", "reason": "model_ended" }
  ```
</ResponseField>

<ResponseField name="error" type="non-fatal error">
  A server-side error report. The call continues when possible.

  ```json theme={null}
  { "type": "error", "data": "..." }
  ```
</ResponseField>

## Call diagnostics

Turn-taking lives on the server and in the voice model, so a client that hears silence cannot
tell why. The `debug` message opens a second stream of frames that says what the server sees.
It is meant for the in-app call debugger, not for product features: the shape of these frames can
change without notice.

<ResponseField name="debug_backlog" type="sent once, on enable">
  Every diagnostic event the server recorded since the call started (the last 300). Enabling the
  stream after a problem still shows what led up to it.

  ```json theme={null}
  { "type": "debug_backlog", "events": [ { "t": 41230, "ev": "vad_end", "speechMs": 9100 } ] }
  ```
</ResponseField>

<ResponseField name="debug" type="one event">
  One diagnostic event. `t` is milliseconds since the call started, on the server clock. `ev`
  names the event; the other fields depend on it. Transcript events carry character counts,
  never text.

  ```json theme={null}
  { "type": "debug", "t": 46231, "ev": "no_reply", "quietMs": 5001, "speechMs": 9100, "transcriptChars": 212, "micGapMs": 64 }
  ```
</ResponseField>

<ResponseField name="debug_state" type="every second">
  A snapshot of the server side of the call: the inferred turn state, the server's own energy
  VAD over the received microphone audio, how long ago the last microphone chunk arrived, the
  state of the upstream model socket, and the result of the search for coaching examples.

  ```json theme={null}
  {
    "type": "debug_state", "t": 47000,
    "turn": "awaiting_reply", "turnSinceMs": 5800,
    "vad": { "speaking": false, "levelDb": -58.2, "floorDb": -61.0, "runMs": 6600, "totalMs": 47000 },
    "micSinceMs": 60, "micChunksPerS": 15, "clientBufferedBytes": 0,
    "gemini": { "socket": "open", "ready": true, "pendingAudio": 0, "reconnects": 1, "hasResumeHandle": true, "sinceFrameMs": 5900, "sinceAudioSentMs": 60, "bufferedBytes": 0, "rollPending": false },
    "rag": { "outcome": "ok", "chunks": 2, "tookMs": 412 }
  }
  ```
</ResponseField>

This is a `rag` event:

```json theme={null}
{
  "type": "debug", "t": 3, "ev": "rag", "outcome": "ok", "tookMs": 412, "queryChars": 380,
  "topK": 2, "minSimilarity": 0.55, "maxChars": 1500,
  "candidates": [
    { "id": "7c1e...", "datasetId": "b2a0...", "similarity": 0.712, "chars": 640, "verdict": "used" },
    { "id": "90af...", "datasetId": "b2a0...", "similarity": 0.498, "chars": 580, "verdict": "weak" }
  ]
}
```

The events, in the order a normal exchange produces them:

| `ev` | Meaning |
| - | - |
| `rag` | The first event of each call, at `t` near 0. It says what the search for coaching examples did. `outcome` is `ok`, `empty` (no example passed the limits), `cached` (the query did not change since the last search for this student, so the server used the kept result and did no search), `stale` (the search was late or failed, and the server used the kept result of an older query), `timeout`, `error`, `no_query` (the student has no chat history) or `off` (`rag=off`). `tookMs` is the time of the search. `embedMs` and `searchMs` divide that time into the embedding request and the database search, when the search completed. A `searchMs` above 1500 usually means that the vector index is missing. `cacheAgeMs` is the age of the kept result, when the server used one. A search that is late continues in the background, and the server keeps its result for the next call (24 hours, per server replica). `candidates` lists each example that the search found, with `similarity`, `chars` and a `verdict`: `used`, `weak` (below `minSimilarity`), `trimmed` (the example was cut at the end of a line to fit in `maxChars`, and `usedChars` is its new length), `too_long` (less than 400 characters of `maxChars` remained) or `over_top_k`. The event carries ids and numbers. `content` says if the text of the examples is included: when it is `true`, each candidate has a `content` field (at most 2000 characters). The server sends the text to all clients at this time, and can limit it to a list of testers without an app release. |
| `rag_tool` | The model called its `koc_ornekleri` tool during the call, to search for coaching examples on the topic that the student started. The payload is the same as `rag`. This search has a limit of 2000 ms and a budget of 1600 characters. When the search is late or fails, the tool returns an empty list and Lala continues without examples. `rag=off` and the server setting `CALL_RAG_TOOL=false` remove the tool. |
| `client_speech_end` | The client sent `speech_end` and the server forwarded it to the model. The reply clock starts here, not at the later `vad_end`. |
| `vad_start`, `vad_end` | The server's energy VAD heard the student start and stop. `vad_end` carries `speechMs` and `transcriptChars` (how much the model transcribed during the burst). |
| `turn` | The inferred turn state changed: `listening`, `user_speaking`, `awaiting_reply`, `lala_speaking`. |
| `transcript` | A transcription delta arrived (`speaker`, `chars`, `final`). |
| `reply_start` | The first byte of Lala's reply, with `afterQuietMs` since the server VAD closed the student's turn. |
| `turn_complete`, `interrupted` | Lala's turn ended, with `voiceMs` of audio sent for it. |
| `empty_turn` | The model closed a turn without any audio. The student hears nothing. |
| `no_reply` | The student has been quiet for 5 seconds and no reply started. `transcriptChars` near zero means the model never heard speech; `micGapMs` large means the server stopped receiving audio, so no silence reached the model and its VAD never closed the turn. |
| `gemini_ready`, `gemini_close`, `gemini_reconnect`, `gemini_reconnect_failed`, `gemini_gone` | Lifecycle of the upstream model socket. |
| `gemini_go_away`, `gemini_roll` | The model announced the end of its \~10-minute connection; the server rolled onto a new one at the next quiet moment (`trigger: idle`) or at the deadline. |

`no_reply` and `empty_turn` are also written to the server log for every call, with or without the
stream enabled.

## Reconnect after a network drop

The call survives a short network drop, for example a switch from WiFi to cellular. When the
socket drops without an `ended` message, the server keeps the call alive for 25 seconds.

1. Open a new socket to the same URL, with the `callId` from the `ready` message added:

```text theme={null}
wss://client-api.lala.ist/api/call/ws?token=<jwt>&callId=<callId>
```

2. Wait for `ready`. On a resumed call it contains `"resumed": true`. Do not send `context`
   again — the prompt is set at call start.
3. Continue to send `audio` messages.

Microphone audio that you send in the seconds around the drop is not lost. The server holds it
and replays it to the model.

If the server closes with code `4004`, the grace period is over or a different server instance
answered. Start a new call in that case.

The server also sends a WebSocket protocol ping every 20 seconds. WebSocket libraries answer
pings automatically — you do not have to write code for this. A client that does not answer two
pings is disconnected and gets the 25-second reconnect window.

## After the call

The server stores the full transcript with the call record. It also writes a summary into Lala's
conversation memory. As a result, the next chat message knows what was said on the call. The
client does not have to send anything for this.

## Limits

* A call stops after 30 minutes.
* A network drop longer than 25 seconds ends the call.
* One short reconnect gap is possible mid-call while the server renews its model session. The
  socket to the client stays open through it, and microphone audio from the gap is replayed.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.