Guides

Streaming

How streamed responses are delivered, what each chunk contains and how to read a stream.

View as Markdown

Some endpoints can send their result while it is being produced instead of all at once: a chat answer when the request sets "stream": true, and a workflow run. All streams use the same envelope, described here; what the envelope carries depends on the endpoint. This page describes the envelope and the chat stream; the workflow run events are described with the workflow endpoint.

How a stream is delivered

A streamed response is an ordinary HTTP response whose body arrives piece by piece:

  • Status 200, Content-Type: text/event-stream; charset=utf-8, chunked transfer encoding (no Content-Length).
  • The body is a sequence of chunks. Each chunk is one JSON object, and the objects are written directly one after another, with no separator: no newline, no data: prefix.
text
{"timestamp": 1759406400810, "data": "Unit tests", "error": null}{"timestamp": 1759406400834, "data": " catch", "error": null}{"timestamp": 1759406400851, "data": " regressions", "error": null}

Despite the content type, the body is not in the Server-Sent Events format, so EventSource and SSE client libraries cannot read it (and EventSource cannot send a POST with a body anyway). Read the body as a byte stream and split it into JSON objects yourself, see Reading a stream. A network read may end in the middle of an object or contain several objects, so always buffer.

Errors that are detected before the stream starts (authentication, permissions, invalid request, unknown model, ...) are returned as a normal error response with their own HTTP status and a {"detail": "..."} body, see Errors. Check the status code before reading the body as a stream.

Chunks

Field Type Description
timestamp timestamp When the chunk was produced (milliseconds since the epoch).
data string, object or null The payload. Its shape depends on the endpoint and, for chat, on runner_version. null in an error chunk.
error Stream error or null Set when the request failed after the stream started; null otherwise.

Keepalive. When nothing has been produced for about a second (the model is thinking, a tool is running), the server sends a chunk with data set to the empty string "". Ignore these chunks; they only keep the connection and any proxies from timing out. Keepalives are sent in every stream format, so data can be "" even in streams whose payloads are otherwise objects.

Errors in a stream

Once the stream has started, the HTTP status is already 200. A failure after that point is sent as a last chunk with error set, and the stream ends.

json
{"timestamp": 1759406402117, "data": null, "error": {"status": 422, "text": "Validation error", "detail": "not enough user credit"}}
Field Type Description
status integer The HTTP status the error would have had as a normal response, for example 422 or 500.
text string Short name of the error class: Validation error, Forbidden, Not found, Unauthorized, Too many requests, Lock error, Internal server error; Internal Server Error for unexpected failures.
detail string The error message, the same text a normal error response carries in detail.

The chat errors that can arrive this way are marked during in the chat error table; the most common one is running out of credits in the middle of an answer. A chat answer interrupted by an error is kept in the conversation up to the point where it stopped.

End of the stream

The stream ends when the HTTP response body ends. There is no final "end" chunk:

  • the body ended after a chunk with error set: the request failed;
  • the body ended otherwise: the request completed.

In the chat event stream, a done event marks the moment the model finished its answer, but treat the end of the body as the end of the stream.

If your client closes the connection early, the server stops generating the chat answer. The part generated so far is kept in the conversation and model calls that already ran are charged. (A workflow run continues on the server after a disconnect.)

Chat stream formats

The chat request field runner_version selects what data contains in a chat stream:

runner_version data Use it when
"1" (default) A string: the next piece of the answer text. You only need the answer text.
"2" An event object: text, reasoning, tool activity, status. You want to show progress, reasoning or tool activity separately from the text.

Both formats come from the same generation and the stored answer is the same; only what is sent over the wire differs.

A model without the stream capability (see Models and tools) can still be streamed: its answer text arrives in one piece (one text fragment, or one text event) when it is complete, and it is stored like any other answer.

Text stream (version 1)

Each data is a fragment of the answer as Markdown. Concatenate the fragments in order to get the full answer; keepalive chunks ("") add nothing.

Tool calls are part of the text, exactly as in the content of a non-streamed answer: each tool call is a block that starts with the marker "\n\n### calling AI tool ...\n", continues with the progress text the tool reports and ends with the marker "\n***\n". See Tool calls in the answer. Reasoning, status information and the internal steps of sub-assistants are not sent in this format.

text
{"timestamp": 1759406400120, "data": "", "error": null}{"timestamp": 1759406400810, "data": "Let me check", "error": null}{"timestamp": 1759406400833, "data": " that.", "error": null}{"timestamp": 1759406401002, "data": "\n\n### calling AI tool ...\n", "error": null}{"timestamp": 1759406401004, "data": "### Scraping URL: https://example.com/rates\n", "error": null}{"timestamp": 1759406402950, "data": "\n***\n", "error": null}{"timestamp": 1759406403410, "data": "The current rate is 24.35 CZK per euro.", "error": null}

Event stream (version 2)

Each data (other than keepalives) is an event object. All fields are always present; the ones that do not apply to the event type are empty ("", null or {}).

Field Type Description
type string Event type, see Event types.
delta string Text carried by text, reasoning and tool_call_output events; "" otherwise.
phase string or null Phase of a status event, see Status phases.
call_id string or null Id of the tool call the event belongs to (tool events, nested, some status events).
tool_name string or null Technical name of the tool, the same id that Models and tools lists and assistants reference.
tool_display_name string or null Human-readable name of the tool.
tool_icon string or null Icon identifier the AYETO app uses for the tool; may be empty.
tool_status string or null Outcome in tool_call_end: ok or error.
data object Extra values of some events (model_selected status, done). {} otherwise.
chunk event or null The inner event of a nested event.

A full chunk of the event stream looks like this:

json
{"timestamp": 1759406400810, "data": {"type": "text", "delta": "Unit tests", "phase": null, "call_id": null, "tool_name": null, "tool_display_name": null, "tool_icon": null, "tool_status": null, "data": {}, "chunk": null}, "error": null}

Event types

type Meaning Fields used
text Next piece of the visible answer text (Markdown). delta
reasoning Next piece of the model's reasoning or thinking summary, for models that expose it. Not part of the answer text. delta
status The request entered a new phase. phase, and depending on the phase call_id, tool_name, tool_display_name, tool_icon, data
tool_call_start A tool started running. call_id, tool_name, tool_display_name, tool_icon
tool_call_output Progress text reported by the running tool (Markdown), for example search queries or a link to a generated file. call_id, delta
tool_call_end The tool finished. call_id, tool_status
nested An event of a sub-assistant that the assistant called as a tool. call_id is the id of that tool call, chunk is the sub-assistant's own event (any type, including nested again for deeper levels). call_id, chunk
done The model finished its answer. data.stop_reason is the provider's reason, for example stop, end_turn or completed, or null. data

Status phases

phase Meaning
context_building The request is being prepared: instructions, tools, history. Usually the first event.
rag_query The assistant's knowledge base is being searched.
relevant_history The relevant part of the history is being selected (relevant_history in the request).
model_selected Names the model that answers: data.model (model id), data.model_name (display name) and, when the assistant uses an automatic model, data.auto_model (id of the automatic model that picked it).
model_call A request was sent to the model and its first tokens are awaited. Sent again after every round of tool calls.
context_compacted The conversation history was shortened to fit the model's context window.
tool_call_pending The model is writing a tool call; call_id and tool_name identify it. The tool starts with tool_call_start.
silent_tool A tool that has no visible output is running; call_id, tool_name, tool_display_name and tool_icon identify it. No tool_call_* events follow for it.

New event types, phases and data keys may be added. Ignore values you do not know.

Order of events

A typical answer that uses one tool produces (keepalives omitted, one data per line):

json
{"type": "status", "phase": "context_building", ...}
{"type": "status", "phase": "model_selected", "data": {"model": "gpt-5-mini", "model_name": "GPT-5 mini"}, ...}
{"type": "status", "phase": "model_call", ...}
{"type": "text", "delta": "Let me check", ...}
{"type": "text", "delta": " that.", ...}
{"type": "status", "phase": "tool_call_pending", "call_id": "call_9f2c", "tool_name": "WebScraperFunctionTool", ...}
{"type": "tool_call_start", "call_id": "call_9f2c", "tool_name": "WebScraperFunctionTool", "tool_display_name": "Web Scraper", "tool_icon": "ayeto", ...}
{"type": "tool_call_output", "call_id": "call_9f2c", "delta": "### Scraping URL: https://example.com/rates\n", ...}
{"type": "tool_call_end", "call_id": "call_9f2c", "tool_status": "ok", ...}
{"type": "status", "phase": "model_call", ...}
{"type": "text", "delta": "The current rate is 24.35 CZK per euro.", ...}
{"type": "done", "data": {"stop_reason": "stop"}, ...}

Rebuilding the answer text

To get the same text a text stream or the stored message contains, append for each top-level event:

Event Append
text delta
tool_call_start "\n\n### calling AI tool ...\n"
tool_call_output delta
tool_call_end "\n***\n"
any other nothing

Generated files (for example images) appear as Markdown links in tool_call_output. After the stream ends you can read the stored answer, including its attachments, through the Conversations API.

Reading a stream

The examples send a chat message with runner_version "2", print the answer text as it arrives and stop on an error chunk. Both split the body into JSON objects with a small buffer, so they work however the network splits the data.

Python

python
import codecs
import json
import os
import uuid

import requests


def iter_chunks(response):
    """Yield the JSON chunks of a streamed response."""
    decoder = json.JSONDecoder()
    text = codecs.getincrementaldecoder("utf-8")()
    buffer = ""
    for raw in response.iter_content(chunk_size=None):
        buffer += text.decode(raw)
        while True:
            buffer = buffer.lstrip()
            if not buffer:
                break
            try:
                chunk, end = decoder.raw_decode(buffer)
            except json.JSONDecodeError:
                break  # incomplete object, wait for more data
            buffer = buffer[end:]
            yield chunk


response = requests.post(
    "https://ayeto.ai/api/v3/chat",
    headers={"uni-api-key": os.environ["AYETO_API_KEY"]},
    json={
        "conversation_id": str(uuid.uuid4()),
        "model": "gpt-5-mini",
        "message": "Explain recursion with a short example.",
        "stream": True,
        "runner_version": "2",
    },
    stream=True,
    timeout=(10, 600),
)
if response.status_code != 200:  # failed before the stream started
    raise RuntimeError(f"{response.status_code}: {response.json()['detail']}")

for chunk in iter_chunks(response):
    if chunk["error"]:
        error = chunk["error"]
        raise RuntimeError(f"{error['status']}: {error['detail']}")
    event = chunk["data"]
    if not isinstance(event, dict):
        continue  # keepalive
    if event["type"] == "text":
        print(event["delta"], end="", flush=True)
    elif event["type"] == "tool_call_start":
        print(f"\n[{event['tool_display_name']}]", flush=True)
    elif event["type"] == "status" and event["phase"] == "model_selected":
        print(f"[answering: {event['data']['model']}]", flush=True)
print()

With runner_version "1", chunk["data"] is a string: print or append it as it is.

JavaScript

javascript
// Yields the JSON chunks of a streamed response.
async function* readChunks(response) {
  const reader = response.body.getReader();
  const decoder = new TextDecoder();
  let buffer = "";
  let pos = 0;
  let depth = 0;
  let inString = false;
  let escaped = false;
  while (true) {
    const { value, done } = await reader.read();
    if (done) break;
    buffer += decoder.decode(value, { stream: true });
    for (; pos < buffer.length; pos++) {
      const c = buffer[pos];
      if (inString) {
        if (escaped) escaped = false;
        else if (c === "\\") escaped = true;
        else if (c === '"') inString = false;
      } else if (c === '"') {
        inString = true;
      } else if (c === "{") {
        depth++;
      } else if (c === "}" && --depth === 0) {
        yield JSON.parse(buffer.slice(0, pos + 1));
        buffer = buffer.slice(pos + 1);
        pos = -1; // continue at the start of the remaining buffer
      }
    }
  }
}

const response = await fetch("https://ayeto.ai/api/v3/chat", {
  method: "POST",
  headers: {
    "uni-api-key": process.env.AYETO_API_KEY,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    conversation_id: crypto.randomUUID(),
    model: "gpt-5-mini",
    message: "Explain recursion with a short example.",
    stream: true,
    runner_version: "2",
  }),
});
if (!response.ok) {
  // failed before the stream started
  const body = await response.json();
  throw new Error(`${response.status}: ${JSON.stringify(body.detail)}`);
}

let answer = "";
for await (const chunk of readChunks(response)) {
  if (chunk.error) {
    throw new Error(`${chunk.error.status}: ${chunk.error.detail}`);
  }
  const event = chunk.data;
  if (typeof event !== "object" || event === null) continue; // keepalive
  if (event.type === "text") {
    answer += event.delta;
    process.stdout.write(event.delta);
  }
}

The same functions read a workflow run stream; only the data payloads differ, see Stream a run.