Some endpoints can send their result while it is being produced instead of all at once:
a chat answer when the request sets "stream": true, and a
workflow run. All streams use the same envelope, described
here; what the envelope carries depends on the endpoint. This page describes the
envelope and the chat stream; the workflow run events are described with the
workflow endpoint.
How a stream is delivered
A streamed response is an ordinary HTTP response whose body arrives piece by piece:
- Status
200,Content-Type: text/event-stream; charset=utf-8, chunked transfer encoding (noContent-Length). - The body is a sequence of chunks. Each chunk is one JSON object, and the
objects are written directly one after another, with no separator: no newline,
no
data:prefix.
{"timestamp": 1759406400810, "data": "Unit tests", "error": null}{"timestamp": 1759406400834, "data": " catch", "error": null}{"timestamp": 1759406400851, "data": " regressions", "error": null}
Despite the content type, the body is not in the Server-Sent Events format, so
EventSource and SSE client libraries cannot read it (and EventSource cannot send a
POST with a body anyway). Read the body as a byte stream and split it into JSON objects
yourself, see Reading a stream. A network read may end in the
middle of an object or contain several objects, so always buffer.
Errors that are detected before the stream starts (authentication, permissions, invalid
request, unknown model, ...) are returned as a normal error response with their own HTTP
status and a {"detail": "..."} body, see Errors. Check the
status code before reading the body as a stream.
Chunks
| Field | Type | Description |
|---|---|---|
timestamp |
timestamp | When the chunk was produced (milliseconds since the epoch). |
data |
string, object or null | The payload. Its shape depends on the endpoint and, for chat, on runner_version. null in an error chunk. |
error |
Stream error or null | Set when the request failed after the stream started; null otherwise. |
Keepalive. When nothing has been produced for about a second (the model is thinking,
a tool is running), the server sends a chunk with data set to the empty string "".
Ignore these chunks; they only keep the connection and any proxies from timing out.
Keepalives are sent in every stream format, so data can be "" even in streams whose
payloads are otherwise objects.
Errors in a stream
Once the stream has started, the HTTP status is already 200. A failure after that point
is sent as a last chunk with error set, and the stream ends.
{"timestamp": 1759406402117, "data": null, "error": {"status": 422, "text": "Validation error", "detail": "not enough user credit"}}
| Field | Type | Description |
|---|---|---|
status |
integer | The HTTP status the error would have had as a normal response, for example 422 or 500. |
text |
string | Short name of the error class: Validation error, Forbidden, Not found, Unauthorized, Too many requests, Lock error, Internal server error; Internal Server Error for unexpected failures. |
detail |
string | The error message, the same text a normal error response carries in detail. |
The chat errors that can arrive this way are marked during in the chat error table; the most common one is running out of credits in the middle of an answer. A chat answer interrupted by an error is kept in the conversation up to the point where it stopped.
End of the stream
The stream ends when the HTTP response body ends. There is no final "end" chunk:
- the body ended after a chunk with
errorset: the request failed; - the body ended otherwise: the request completed.
In the chat event stream, a done event marks the moment the
model finished its answer, but treat the end of the body as the end of the stream.
If your client closes the connection early, the server stops generating the chat answer. The part generated so far is kept in the conversation and model calls that already ran are charged. (A workflow run continues on the server after a disconnect.)
Chat stream formats
The chat request field runner_version selects what data contains in a chat stream:
runner_version |
data |
Use it when |
|---|---|---|
"1" (default) |
A string: the next piece of the answer text. | You only need the answer text. |
"2" |
An event object: text, reasoning, tool activity, status. | You want to show progress, reasoning or tool activity separately from the text. |
Both formats come from the same generation and the stored answer is the same; only what is sent over the wire differs.
A model without the stream capability (see Models and tools)
can still be streamed: its answer text arrives in one piece (one text fragment, or one
text event) when it is complete, and it is stored like any other answer.
Text stream (version 1)
Each data is a fragment of the answer as Markdown. Concatenate the fragments in order
to get the full answer; keepalive chunks ("") add nothing.
Tool calls are part of the text, exactly as in the content of a non-streamed
answer: each tool call is a block that starts with the marker
"\n\n### calling AI tool ...\n", continues with the progress text the tool reports and
ends with the marker "\n***\n". See Tool calls in the answer.
Reasoning, status information and the internal steps of sub-assistants are not sent in
this format.
{"timestamp": 1759406400120, "data": "", "error": null}{"timestamp": 1759406400810, "data": "Let me check", "error": null}{"timestamp": 1759406400833, "data": " that.", "error": null}{"timestamp": 1759406401002, "data": "\n\n### calling AI tool ...\n", "error": null}{"timestamp": 1759406401004, "data": "### Scraping URL: https://example.com/rates\n", "error": null}{"timestamp": 1759406402950, "data": "\n***\n", "error": null}{"timestamp": 1759406403410, "data": "The current rate is 24.35 CZK per euro.", "error": null}
Event stream (version 2)
Each data (other than keepalives) is an event object. All fields are always present;
the ones that do not apply to the event type are empty ("", null or {}).
| Field | Type | Description |
|---|---|---|
type |
string | Event type, see Event types. |
delta |
string | Text carried by text, reasoning and tool_call_output events; "" otherwise. |
phase |
string or null | Phase of a status event, see Status phases. |
call_id |
string or null | Id of the tool call the event belongs to (tool events, nested, some status events). |
tool_name |
string or null | Technical name of the tool, the same id that Models and tools lists and assistants reference. |
tool_display_name |
string or null | Human-readable name of the tool. |
tool_icon |
string or null | Icon identifier the AYETO app uses for the tool; may be empty. |
tool_status |
string or null | Outcome in tool_call_end: ok or error. |
data |
object | Extra values of some events (model_selected status, done). {} otherwise. |
chunk |
event or null | The inner event of a nested event. |
A full chunk of the event stream looks like this:
{"timestamp": 1759406400810, "data": {"type": "text", "delta": "Unit tests", "phase": null, "call_id": null, "tool_name": null, "tool_display_name": null, "tool_icon": null, "tool_status": null, "data": {}, "chunk": null}, "error": null}
Event types
type |
Meaning | Fields used |
|---|---|---|
text |
Next piece of the visible answer text (Markdown). | delta |
reasoning |
Next piece of the model's reasoning or thinking summary, for models that expose it. Not part of the answer text. | delta |
status |
The request entered a new phase. | phase, and depending on the phase call_id, tool_name, tool_display_name, tool_icon, data |
tool_call_start |
A tool started running. | call_id, tool_name, tool_display_name, tool_icon |
tool_call_output |
Progress text reported by the running tool (Markdown), for example search queries or a link to a generated file. | call_id, delta |
tool_call_end |
The tool finished. | call_id, tool_status |
nested |
An event of a sub-assistant that the assistant called as a tool. call_id is the id of that tool call, chunk is the sub-assistant's own event (any type, including nested again for deeper levels). |
call_id, chunk |
done |
The model finished its answer. data.stop_reason is the provider's reason, for example stop, end_turn or completed, or null. |
data |
Status phases
phase |
Meaning |
|---|---|
context_building |
The request is being prepared: instructions, tools, history. Usually the first event. |
rag_query |
The assistant's knowledge base is being searched. |
relevant_history |
The relevant part of the history is being selected (relevant_history in the request). |
model_selected |
Names the model that answers: data.model (model id), data.model_name (display name) and, when the assistant uses an automatic model, data.auto_model (id of the automatic model that picked it). |
model_call |
A request was sent to the model and its first tokens are awaited. Sent again after every round of tool calls. |
context_compacted |
The conversation history was shortened to fit the model's context window. |
tool_call_pending |
The model is writing a tool call; call_id and tool_name identify it. The tool starts with tool_call_start. |
silent_tool |
A tool that has no visible output is running; call_id, tool_name, tool_display_name and tool_icon identify it. No tool_call_* events follow for it. |
New event types, phases and data keys may be added. Ignore values you do not know.
Order of events
A typical answer that uses one tool produces (keepalives omitted, one data per line):
{"type": "status", "phase": "context_building", ...}
{"type": "status", "phase": "model_selected", "data": {"model": "gpt-5-mini", "model_name": "GPT-5 mini"}, ...}
{"type": "status", "phase": "model_call", ...}
{"type": "text", "delta": "Let me check", ...}
{"type": "text", "delta": " that.", ...}
{"type": "status", "phase": "tool_call_pending", "call_id": "call_9f2c", "tool_name": "WebScraperFunctionTool", ...}
{"type": "tool_call_start", "call_id": "call_9f2c", "tool_name": "WebScraperFunctionTool", "tool_display_name": "Web Scraper", "tool_icon": "ayeto", ...}
{"type": "tool_call_output", "call_id": "call_9f2c", "delta": "### Scraping URL: https://example.com/rates\n", ...}
{"type": "tool_call_end", "call_id": "call_9f2c", "tool_status": "ok", ...}
{"type": "status", "phase": "model_call", ...}
{"type": "text", "delta": "The current rate is 24.35 CZK per euro.", ...}
{"type": "done", "data": {"stop_reason": "stop"}, ...}
Rebuilding the answer text
To get the same text a text stream or the stored message contains, append for each top-level event:
| Event | Append |
|---|---|
text |
delta |
tool_call_start |
"\n\n### calling AI tool ...\n" |
tool_call_output |
delta |
tool_call_end |
"\n***\n" |
| any other | nothing |
Generated files (for example images) appear as Markdown links in tool_call_output.
After the stream ends you can read the stored answer, including its attachments, through
the Conversations API.
Reading a stream
The examples send a chat message with runner_version "2", print the answer text as it
arrives and stop on an error chunk. Both split the body into JSON objects with a small
buffer, so they work however the network splits the data.
Python
import codecs
import json
import os
import uuid
import requests
def iter_chunks(response):
"""Yield the JSON chunks of a streamed response."""
decoder = json.JSONDecoder()
text = codecs.getincrementaldecoder("utf-8")()
buffer = ""
for raw in response.iter_content(chunk_size=None):
buffer += text.decode(raw)
while True:
buffer = buffer.lstrip()
if not buffer:
break
try:
chunk, end = decoder.raw_decode(buffer)
except json.JSONDecodeError:
break # incomplete object, wait for more data
buffer = buffer[end:]
yield chunk
response = requests.post(
"https://ayeto.ai/api/v3/chat",
headers={"uni-api-key": os.environ["AYETO_API_KEY"]},
json={
"conversation_id": str(uuid.uuid4()),
"model": "gpt-5-mini",
"message": "Explain recursion with a short example.",
"stream": True,
"runner_version": "2",
},
stream=True,
timeout=(10, 600),
)
if response.status_code != 200: # failed before the stream started
raise RuntimeError(f"{response.status_code}: {response.json()['detail']}")
for chunk in iter_chunks(response):
if chunk["error"]:
error = chunk["error"]
raise RuntimeError(f"{error['status']}: {error['detail']}")
event = chunk["data"]
if not isinstance(event, dict):
continue # keepalive
if event["type"] == "text":
print(event["delta"], end="", flush=True)
elif event["type"] == "tool_call_start":
print(f"\n[{event['tool_display_name']}]", flush=True)
elif event["type"] == "status" and event["phase"] == "model_selected":
print(f"[answering: {event['data']['model']}]", flush=True)
print()
With runner_version "1", chunk["data"] is a string: print or append it as it is.
JavaScript
// Yields the JSON chunks of a streamed response.
async function* readChunks(response) {
const reader = response.body.getReader();
const decoder = new TextDecoder();
let buffer = "";
let pos = 0;
let depth = 0;
let inString = false;
let escaped = false;
while (true) {
const { value, done } = await reader.read();
if (done) break;
buffer += decoder.decode(value, { stream: true });
for (; pos < buffer.length; pos++) {
const c = buffer[pos];
if (inString) {
if (escaped) escaped = false;
else if (c === "\\") escaped = true;
else if (c === '"') inString = false;
} else if (c === '"') {
inString = true;
} else if (c === "{") {
depth++;
} else if (c === "}" && --depth === 0) {
yield JSON.parse(buffer.slice(0, pos + 1));
buffer = buffer.slice(pos + 1);
pos = -1; // continue at the start of the remaining buffer
}
}
}
}
const response = await fetch("https://ayeto.ai/api/v3/chat", {
method: "POST",
headers: {
"uni-api-key": process.env.AYETO_API_KEY,
"Content-Type": "application/json",
},
body: JSON.stringify({
conversation_id: crypto.randomUUID(),
model: "gpt-5-mini",
message: "Explain recursion with a short example.",
stream: true,
runner_version: "2",
}),
});
if (!response.ok) {
// failed before the stream started
const body = await response.json();
throw new Error(`${response.status}: ${JSON.stringify(body.detail)}`);
}
let answer = "";
for await (const chunk of readChunks(response)) {
if (chunk.error) {
throw new Error(`${chunk.error.status}: ${chunk.error.detail}`);
}
const event = chunk.data;
if (typeof event !== "object" || event === null) continue; // keepalive
if (event.type === "text") {
answer += event.delta;
process.stdout.write(event.delta);
}
}
The same functions read a workflow run stream; only the data payloads differ, see
Stream a run.