# Streaming

> How streamed responses are delivered, what each chunk contains and how to read a stream.

Some endpoints can send their result while it is being produced instead of all at once:
a [chat](chat.md) answer when the request sets `"stream": true`, and a
[workflow run](workflows.md#stream-a-run). All streams use the same envelope, described
here; what the envelope carries depends on the endpoint. This page describes the
envelope and the chat stream; the workflow run events are described with the
[workflow endpoint](workflows.md#stream-a-run).

## How a stream is delivered

A streamed response is an ordinary HTTP response whose body arrives piece by piece:

- Status `200`, `Content-Type: text/event-stream; charset=utf-8`, chunked transfer
  encoding (no `Content-Length`).
- The body is a sequence of [chunks](#chunks). Each chunk is one JSON object, and the
  objects are written **directly one after another, with no separator**: no newline,
  no `data:` prefix.

```text
{"timestamp": 1759406400810, "data": "Unit tests", "error": null}{"timestamp": 1759406400834, "data": " catch", "error": null}{"timestamp": 1759406400851, "data": " regressions", "error": null}
```

Despite the content type, the body is **not** in the Server-Sent Events format, so
`EventSource` and SSE client libraries cannot read it (and `EventSource` cannot send a
`POST` with a body anyway). Read the body as a byte stream and split it into JSON objects
yourself, see [Reading a stream](#reading-a-stream). A network read may end in the
middle of an object or contain several objects, so always buffer.

Errors that are detected before the stream starts (authentication, permissions, invalid
request, unknown model, ...) are returned as a normal error response with their own HTTP
status and a `{"detail": "..."}` body, see [Errors](conventions.md#errors). Check the
status code before reading the body as a stream.

## Chunks

| Field | Type | Description |
|---|---|---|
| `timestamp` | timestamp | When the chunk was produced (milliseconds since the epoch). |
| `data` | string, object or null | The payload. Its shape depends on the endpoint and, for chat, on [`runner_version`](#chat-stream-formats). `null` in an error chunk. |
| `error` | [Stream error](#errors-in-a-stream) or null | Set when the request failed after the stream started; `null` otherwise. |

**Keepalive.** When nothing has been produced for about a second (the model is thinking,
a tool is running), the server sends a chunk with `data` set to the empty string `""`.
Ignore these chunks; they only keep the connection and any proxies from timing out.
Keepalives are sent in every stream format, so `data` can be `""` even in streams whose
payloads are otherwise objects.

## Errors in a stream

Once the stream has started, the HTTP status is already `200`. A failure after that point
is sent as a last chunk with `error` set, and the stream ends.

```json
{"timestamp": 1759406402117, "data": null, "error": {"status": 422, "text": "Validation error", "detail": "not enough user credit"}}
```

| Field | Type | Description |
|---|---|---|
| `status` | integer | The HTTP status the error would have had as a normal response, for example `422` or `500`. |
| `text` | string | Short name of the error class: `Validation error`, `Forbidden`, `Not found`, `Unauthorized`, `Too many requests`, `Lock error`, `Internal server error`; `Internal Server Error` for unexpected failures. |
| `detail` | string | The error message, the same text a normal error response carries in `detail`. |

The chat errors that can arrive this way are marked *during* in the
[chat error table](chat.md#errors); the most common one is running out of credits in the
middle of an answer. A chat answer interrupted by an error is kept in the conversation up
to the point where it stopped.

## End of the stream

The stream ends when the HTTP response body ends. There is no final "end" chunk:

- the body ended after a chunk with `error` set: the request failed;
- the body ended otherwise: the request completed.

In the chat [event stream](#event-stream-version-2), a `done` event marks the moment the
model finished its answer, but treat the end of the body as the end of the stream.

If your client closes the connection early, the server stops generating the chat answer.
The part generated so far is kept in the conversation and model calls that already ran
are charged. (A workflow run continues on the server after a disconnect.)

## Chat stream formats

The chat request field `runner_version` selects what `data` contains in a chat stream:

| `runner_version` | `data` | Use it when |
|---|---|---|
| `"1"` (default) | A string: the next piece of the answer text. | You only need the answer text. |
| `"2"` | An [event](#event-stream-version-2) object: text, reasoning, tool activity, status. | You want to show progress, reasoning or tool activity separately from the text. |

Both formats come from the same generation and the stored answer is the same; only what
is sent over the wire differs.

A model without the `stream` capability (see [Models and tools](models-and-tools.md#capabilities))
can still be streamed: its answer text arrives in one piece (one text fragment, or one
`text` event) when it is complete, and it is stored like any other answer.

### Text stream (version 1)

Each `data` is a fragment of the answer as Markdown. Concatenate the fragments in order
to get the full answer; keepalive chunks (`""`) add nothing.

Tool calls are part of the text, exactly as in the `content` of a non-streamed
answer: each tool call is a block that starts with the marker
`"\n\n### calling AI tool ...\n"`, continues with the progress text the tool reports and
ends with the marker `"\n***\n"`. See [Tool calls in the answer](chat.md#tool-calls-in-the-answer).
Reasoning, status information and the internal steps of sub-assistants are not sent in
this format.

```text
{"timestamp": 1759406400120, "data": "", "error": null}{"timestamp": 1759406400810, "data": "Let me check", "error": null}{"timestamp": 1759406400833, "data": " that.", "error": null}{"timestamp": 1759406401002, "data": "\n\n### calling AI tool ...\n", "error": null}{"timestamp": 1759406401004, "data": "### Scraping URL: https://example.com/rates\n", "error": null}{"timestamp": 1759406402950, "data": "\n***\n", "error": null}{"timestamp": 1759406403410, "data": "The current rate is 24.35 CZK per euro.", "error": null}
```

### Event stream (version 2)

Each `data` (other than keepalives) is an event object. All fields are always present;
the ones that do not apply to the event type are empty (`""`, `null` or `{}`).

| Field | Type | Description |
|---|---|---|
| `type` | string | Event type, see [Event types](#event-types). |
| `delta` | string | Text carried by `text`, `reasoning` and `tool_call_output` events; `""` otherwise. |
| `phase` | string or null | Phase of a `status` event, see [Status phases](#status-phases). |
| `call_id` | string or null | Id of the tool call the event belongs to (tool events, `nested`, some `status` events). |
| `tool_name` | string or null | Technical name of the tool, the same id that [Models and tools](models-and-tools.md) lists and assistants reference. |
| `tool_display_name` | string or null | Human-readable name of the tool. |
| `tool_icon` | string or null | Icon identifier the AYETO app uses for the tool; may be empty. |
| `tool_status` | string or null | Outcome in `tool_call_end`: `ok` or `error`. |
| `data` | object | Extra values of some events (`model_selected` status, `done`). `{}` otherwise. |
| `chunk` | event or null | The inner event of a `nested` event. |

A full chunk of the event stream looks like this:

```json
{"timestamp": 1759406400810, "data": {"type": "text", "delta": "Unit tests", "phase": null, "call_id": null, "tool_name": null, "tool_display_name": null, "tool_icon": null, "tool_status": null, "data": {}, "chunk": null}, "error": null}
```

### Event types

| `type` | Meaning | Fields used |
|---|---|---|
| `text` | Next piece of the visible answer text (Markdown). | `delta` |
| `reasoning` | Next piece of the model's reasoning or thinking summary, for models that expose it. Not part of the answer text. | `delta` |
| `status` | The request entered a new phase. | `phase`, and depending on the phase `call_id`, `tool_name`, `tool_display_name`, `tool_icon`, `data` |
| `tool_call_start` | A tool started running. | `call_id`, `tool_name`, `tool_display_name`, `tool_icon` |
| `tool_call_output` | Progress text reported by the running tool (Markdown), for example search queries or a link to a generated file. | `call_id`, `delta` |
| `tool_call_end` | The tool finished. | `call_id`, `tool_status` |
| `nested` | An event of a sub-assistant that the assistant called as a tool. `call_id` is the id of that tool call, `chunk` is the sub-assistant's own event (any type, including `nested` again for deeper levels). | `call_id`, `chunk` |
| `done` | The model finished its answer. `data.stop_reason` is the provider's reason, for example `stop`, `end_turn` or `completed`, or `null`. | `data` |

### Status phases

| `phase` | Meaning |
|---|---|
| `context_building` | The request is being prepared: instructions, tools, history. Usually the first event. |
| `rag_query` | The assistant's knowledge base is being searched. |
| `relevant_history` | The relevant part of the history is being selected (`relevant_history` in the request). |
| `model_selected` | Names the model that answers: `data.model` (model id), `data.model_name` (display name) and, when the assistant uses an automatic model, `data.auto_model` (id of the automatic model that picked it). |
| `model_call` | A request was sent to the model and its first tokens are awaited. Sent again after every round of tool calls. |
| `context_compacted` | The conversation history was shortened to fit the model's context window. |
| `tool_call_pending` | The model is writing a tool call; `call_id` and `tool_name` identify it. The tool starts with `tool_call_start`. |
| `silent_tool` | A tool that has no visible output is running; `call_id`, `tool_name`, `tool_display_name` and `tool_icon` identify it. No `tool_call_*` events follow for it. |

New event types, phases and `data` keys may be added. Ignore values you do not know.

### Order of events

A typical answer that uses one tool produces (keepalives omitted, one `data` per line):

```json
{"type": "status", "phase": "context_building", ...}
{"type": "status", "phase": "model_selected", "data": {"model": "gpt-5-mini", "model_name": "GPT-5 mini"}, ...}
{"type": "status", "phase": "model_call", ...}
{"type": "text", "delta": "Let me check", ...}
{"type": "text", "delta": " that.", ...}
{"type": "status", "phase": "tool_call_pending", "call_id": "call_9f2c", "tool_name": "WebScraperFunctionTool", ...}
{"type": "tool_call_start", "call_id": "call_9f2c", "tool_name": "WebScraperFunctionTool", "tool_display_name": "Web Scraper", "tool_icon": "ayeto", ...}
{"type": "tool_call_output", "call_id": "call_9f2c", "delta": "### Scraping URL: https://example.com/rates\n", ...}
{"type": "tool_call_end", "call_id": "call_9f2c", "tool_status": "ok", ...}
{"type": "status", "phase": "model_call", ...}
{"type": "text", "delta": "The current rate is 24.35 CZK per euro.", ...}
{"type": "done", "data": {"stop_reason": "stop"}, ...}
```

### Rebuilding the answer text

To get the same text a [text stream](#text-stream-version-1) or the stored message
contains, append for each top-level event:

| Event | Append |
|---|---|
| `text` | `delta` |
| `tool_call_start` | `"\n\n### calling AI tool ...\n"` |
| `tool_call_output` | `delta` |
| `tool_call_end` | `"\n***\n"` |
| any other | nothing |

Generated files (for example images) appear as Markdown links in `tool_call_output`.
After the stream ends you can read the stored answer, including its attachments, through
the [Conversations](conversations.md) API.

## Reading a stream

The examples send a chat message with `runner_version` `"2"`, print the answer text as it
arrives and stop on an error chunk. Both split the body into JSON objects with a small
buffer, so they work however the network splits the data.

### Python

```python
import codecs
import json
import os
import uuid

import requests


def iter_chunks(response):
    """Yield the JSON chunks of a streamed response."""
    decoder = json.JSONDecoder()
    text = codecs.getincrementaldecoder("utf-8")()
    buffer = ""
    for raw in response.iter_content(chunk_size=None):
        buffer += text.decode(raw)
        while True:
            buffer = buffer.lstrip()
            if not buffer:
                break
            try:
                chunk, end = decoder.raw_decode(buffer)
            except json.JSONDecodeError:
                break  # incomplete object, wait for more data
            buffer = buffer[end:]
            yield chunk


response = requests.post(
    "https://ayeto.ai/api/v3/chat",
    headers={"uni-api-key": os.environ["AYETO_API_KEY"]},
    json={
        "conversation_id": str(uuid.uuid4()),
        "model": "gpt-5-mini",
        "message": "Explain recursion with a short example.",
        "stream": True,
        "runner_version": "2",
    },
    stream=True,
    timeout=(10, 600),
)
if response.status_code != 200:  # failed before the stream started
    raise RuntimeError(f"{response.status_code}: {response.json()['detail']}")

for chunk in iter_chunks(response):
    if chunk["error"]:
        error = chunk["error"]
        raise RuntimeError(f"{error['status']}: {error['detail']}")
    event = chunk["data"]
    if not isinstance(event, dict):
        continue  # keepalive
    if event["type"] == "text":
        print(event["delta"], end="", flush=True)
    elif event["type"] == "tool_call_start":
        print(f"\n[{event['tool_display_name']}]", flush=True)
    elif event["type"] == "status" and event["phase"] == "model_selected":
        print(f"[answering: {event['data']['model']}]", flush=True)
print()
```

With `runner_version` `"1"`, `chunk["data"]` is a string: print or append it as it is.

### JavaScript

```javascript
// Yields the JSON chunks of a streamed response.
async function* readChunks(response) {
  const reader = response.body.getReader();
  const decoder = new TextDecoder();
  let buffer = "";
  let pos = 0;
  let depth = 0;
  let inString = false;
  let escaped = false;
  while (true) {
    const { value, done } = await reader.read();
    if (done) break;
    buffer += decoder.decode(value, { stream: true });
    for (; pos < buffer.length; pos++) {
      const c = buffer[pos];
      if (inString) {
        if (escaped) escaped = false;
        else if (c === "\\") escaped = true;
        else if (c === '"') inString = false;
      } else if (c === '"') {
        inString = true;
      } else if (c === "{") {
        depth++;
      } else if (c === "}" && --depth === 0) {
        yield JSON.parse(buffer.slice(0, pos + 1));
        buffer = buffer.slice(pos + 1);
        pos = -1; // continue at the start of the remaining buffer
      }
    }
  }
}

const response = await fetch("https://ayeto.ai/api/v3/chat", {
  method: "POST",
  headers: {
    "uni-api-key": process.env.AYETO_API_KEY,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    conversation_id: crypto.randomUUID(),
    model: "gpt-5-mini",
    message: "Explain recursion with a short example.",
    stream: true,
    runner_version: "2",
  }),
});
if (!response.ok) {
  // failed before the stream started
  const body = await response.json();
  throw new Error(`${response.status}: ${JSON.stringify(body.detail)}`);
}

let answer = "";
for await (const chunk of readChunks(response)) {
  if (chunk.error) {
    throw new Error(`${chunk.error.status}: ${chunk.error.detail}`);
  }
  const event = chunk.data;
  if (typeof event !== "object" || event === null) continue; // keepalive
  if (event.type === "text") {
    answer += event.delta;
    process.stdout.write(event.delta);
  }
}
```

The same functions read a workflow run stream; only the `data` payloads differ, see
[Stream a run](workflows.md#stream-a-run).
