# Chat

> Send a message to a model or an assistant and get the answer, in one piece or as a stream.

One endpoint covers chat: you send a user message, AYETO runs the model or assistant
(including any AI tools it uses, such as web search or image generation), stores both
messages in a conversation and returns the answer. The answer comes back as one
message object, or as a [stream](streaming.md) when you ask for it.

## Send a message

| | |
|---|---|
| Endpoint | `POST /api/v3/chat` |
| Scope | `ayeto.chat` |
| Rate limit | [default](conventions.md#rate-limits) |
| Response | [Message](#response), or a [stream](streaming.md) when `stream` is `true` |

### Conversations

Every message belongs to a conversation, identified by `conversation_id`. **You create
the id**: generate a UUID (version 4) on your side and send it with the first message.
If no conversation with that id exists yet, AYETO creates it; if it exists and belongs
to the key's user, the message is added to it and the model sees the earlier messages.

The response does not contain the conversation id. If you omit `conversation_id`, a
new conversation is created for that one message and you have no way to continue it,
so always send an id you keep.

A conversation is bound to what it was created with:

- **Assistant.** A conversation created with `assistant_id` is answered by that
  assistant (its instructions, tools, knowledge and model) for its whole life. Sending a
  different `assistant_id` or a `model` later does not switch it to another assistant.
  A conversation created with a `model` stays a plain model conversation.
- **Organization.** A conversation created in an organization stays there. A follow-up
  message that names no organization runs in the conversation's organization; naming
  another organization is refused (see [Organizations and credits](#organizations-and-credits)).

To list, read or delete conversations, use the [Conversations](conversations.md) API.

### Request headers

| Header | Required | Description |
|---|---|---|
| `uni-api-key` | yes | Your [API key](authentication.md). |
| `Content-Type` | yes | `application/json` |
| `language` | no | The user's language code, for example `EN`, `CZ` or `FR`. The model is told the language with every message, so it can answer in it. Default `EN`. |

### Request

| Field | Type | Required | Description |
|---|---|---|---|
| `message` | string | yes | The user's message, Markdown or plain text. A message may activate an assistant skill with `/skill-slug` (see below). |
| `conversation_id` | UUID | no | Conversation to add the message to; created when it does not exist. Generate it yourself, see [Conversations](#conversations). Default: a new random id. |
| `assistant_id` | UUID | one of `assistant_id`, `model` | Assistant to talk to. Must be your own assistant or one shared with the key's user. When set, `model` is ignored and the assistant's model answers. |
| `model` | string | one of `assistant_id`, `model` | Model id to talk to directly, without an assistant, for example `gpt-5-mini`. The ids are listed by [Models and tools](models-and-tools.md). Required when `assistant_id` is not set. |
| `organization_id` | UUID | no | Run the message in this organization: the organization's credits pay for it and its files count against the organization's storage. The key's user must be an admin or member (not a guest) of the organization. Default: the assistant's organization when the assistant belongs to one, otherwise the organization of the conversation when it continues one created in an organization, otherwise none (personal account). |
| `stream` | boolean | no | `true` returns the answer as a [stream](streaming.md) while it is generated. Default `false`. |
| `runner_version` | string | no | Format of the stream: `"1"` streams plain text, `"2"` streams structured events (text, reasoning, tool activity, status). Only used when `stream` is `true`. Default `"1"`. See [Streaming](streaming.md#chat-stream-formats). |
| `attachments` | array of [Attachment](#attachments) | no | Files sent with the message (documents, images). |
| `max_tokens` | integer | no | Upper limit of tokens the model may generate per model call. Must be between 1 and the model's own maximum. Default: the model's default. |
| `dynamic_tools` | boolean | no | Model conversations only: give the model the standard set of AI tools (web search, image generation, document writers and more). Ignored for assistants, whose tools are configured on the assistant. Default `false`. |
| `use_vision` | boolean | no | Let the model look at attached images (models with vision). With `false`, the model is told about attached images but cannot look at them. Default `true`. |
| `relevant_history` | boolean | no | Send the model only the part of the conversation history that is relevant to the new message, selected by an additional model call. Applies only to models that support it and conversations with at least 3 messages; ignored for reasoning models that use tools. Default `false`. |
| `remove_tool_calls` | boolean | no | Remove the [tool call blocks](#tool-calls-in-the-answer) from `content` of the response, leaving only the model's own text. Not applied to streams. Default `false`. |
| `follow_options` | boolean | no | Ask the model to end its answer, when it fits, with a fenced code block tagged `ayeto-follow-options`: a JSON object with an optional `question`, optional `multi_select` and a list of `options` (`label`, `prompt`, optional `description`) the user can pick as the next message. Only useful if your client renders it. Default `false`. |

**Assistant skills.** A word of the form `/slug` (lowercase letters, digits and
hyphens, separated by whitespace) in `message` activates the assistant skill with that
slug, when the key's user owns it or it is shared with them and it belongs to the same
organization context as the request. The skill's instruction and tools apply to that
message. Unknown slugs are left as plain text.

### Attachments

Each attachment is a file encoded into the request body.

| Field | Type | Required | Description |
|---|---|---|---|
| `filename` | string | yes | File name with extension, for example `report.pdf`. |
| `data` | string | yes | Base64 encoded content as a data URI, for example `data:application/pdf;base64,JVBERi0xLjcK...`. Plain base64 without the `data:...;base64,` prefix is accepted too. Whitespace such as line breaks in the base64 text is ignored; any other character outside the base64 alphabet, or wrong padding, refuses the request with `422`. |
| `mime_type` | string | yes | MIME type, for example `application/pdf` or `image/png`. For images the type is detected from the content (JPEG, PNG, WebP) and corrected if it differs. |

The stored size is taken from the decoded content. A `size` field sent by older clients
is ignored.

Attachments are stored as files of the key's user (or of the organization, see
`organization_id`) and kept with the conversation; deleting the conversation deletes them.

- **Size limit.** Each file is limited by the deployment's upload limit, 50 MB by default.
  A larger file is refused with `422 file is too large`.
- **Storage quota.** Attachments count against the storage quota of the user or the
  organization. When it is full the request fails with `422 storage quota exceeded` or
  `422 organization storage quota exceeded`.
- **Checked first, all or nothing.** Attachments are decoded and stored before anything
  else of the request is stored. When one is refused (invalid base64, too large, over the
  storage quota), the request fails with `422`, the attachments of the same request that
  were already stored are deleted again, and the conversation is left as it was: the user
  message is not added and a new conversation is not created. Fix the attachment and send
  the message again with the same `conversation_id`.
- **How the model reads them.** The model is told which files are attached and reads
  them with built-in tools: documents as extracted text, images through vision (when the
  model supports it and `use_vision` is `true`; large images are scaled down for the
  model). Attachments are therefore useful only with models that support AI tools. Text
  extraction covers common office, PDF and text formats, see
  [Files and media](files-and-media.md).

### Response

Without `stream`, the response is the assistant's answer to the message.

| Field | Type | Description |
|---|---|---|
| `id` | UUID | Id of the answer message within the conversation. |
| `timestamp` | timestamp | When the answer message was created (milliseconds since the epoch). |
| `role` | string | `assistant`. |
| `content` | string | The answer as Markdown. Includes [tool call blocks](#tool-calls-in-the-answer) unless `remove_tool_calls` is `true`. |
| `reasoning_content` | string or null | The model's reasoning or thinking summary, for models that expose it; empty or `null` otherwise. |
| `model` | string or null | Id of the model that produced the answer. For an assistant that uses an automatic model, this is the model actually chosen for the message. |
| `attachments` | array of [Answer attachment](#answer-attachments) | Files produced for the user while answering, such as generated images. |
| `tool_runs` | object | Answers of sub-assistants called as tools, see [Sub-assistant runs](#sub-assistant-runs). Empty object when there are none. |

### Answer attachments

| Field | Type | Description |
|---|---|---|
| `filename` | string | File name. |
| `content_type` | string | MIME type, for example `image/png`. |
| `file_url` | string | Link that downloads the file. It works without an API key, so treat it as a secret. |
| `name` | string or null | Optional display name. |
| `description` | string or null | Optional description. |

### Sub-assistant runs

An assistant can call other assistants as tools. What such a sub-assistant answered is
returned in `tool_runs`, keyed by the position of the tool call among the tool call
blocks of `content`, as a string: `"0"` is the first tool call, `"1"` the second.
Only tool calls answered by a sub-assistant have an entry.

| Field | Type | Description |
|---|---|---|
| `content` | string | The sub-assistant's answer, with its own tool call blocks. |
| `tool_runs` | object | Runs inside the sub-assistant's own tool calls, with the same structure. |

### Tool calls in the answer

When the model uses AI tools while answering, each tool call is written into `content`
as a block between two fixed markers, followed by the model's further text:

```text
Let me check the current exchange rate.

### calling AI tool ...
### Scraping URL: https://example.com/rates
Looking up today's EUR/CZK rate
Successfully scraped 1 page(s)

***
The current rate is 24.35 CZK per euro.
```

The opening marker is the exact string `"\n\n### calling AI tool ...\n"` and the closing
marker is `"\n***\n"`. Between them is the progress text the tool reported (Markdown).
Set `remove_tool_calls` to `true` to get `content` without these blocks, or use the
structured [event stream](streaming.md#event-stream-version-2), which reports tool
activity as separate events.

### Generated images and files

Tools that create files for the user (image generation, document writers) report the
file in two ways:

- the tool call block in `content` contains a Markdown link, for an image
  `![image.png](https://ayeto.ai/api/v1/public/file/read?id=...&secret=...)`;
- generated images are listed in `attachments` of the answer.

Image generation is available when the assistant has an image generation tool, or in a
model conversation with `dynamic_tools` set to `true`. Ask for an image in `message`;
there is no separate image flag.

### Organizations and credits

Every model call is charged in credits, to the key's user or, when the request runs in
an organization, to the user's credit in that organization. A request runs in an
organization when `organization_id` is set, when the assistant belongs to an
organization, or when it continues a conversation created in an organization.

A conversation created in an organization stays in it. A follow-up message sent without
`organization_id` (and without an assistant of another organization) runs in the
conversation's organization: it is charged there, its attachments count against the
organization's storage, and the key's user must still be an admin or member of it. A
follow-up that names a different organization, directly or through `assistant_id`, is
refused with `422 the conversation belongs to another organization` before anything is
stored or charged. A conversation created without an organization has no such rule.

Before each model call AYETO estimates its cost and refuses the call when the balance
does not cover it. While an answer is being generated, it is cut off when its cost would
exceed the balance. In both cases the error is `422` with the detail
`not enough user credit` or `not enough organization credit`. Model calls that already
ran are charged, and an answer cut off in the middle is kept in the conversation.

See [Account](account.md) for reading the credit balance.

### Errors

Errors are returned as `{"detail": "..."}` with the HTTP status below. With `stream` set
to `true`, errors marked *during* happen after the stream has started: the HTTP status is
`200` and the error arrives as an [error chunk](streaming.md#errors-in-a-stream) with the
same status and detail. Errors marked *before* are returned as a normal error response in
both modes.

| Status | `detail` | When | Cause |
|---|---|---|---|
| `401` | `API key is invalid` | before | See [Authentication](authentication.md#authentication-errors). |
| `403` | `permission denied` | before | The key's user is not allowed to use chat, or may not use the assistant. |
| `403` | `not shared with user` | before | The assistant is neither the user's own nor shared with them. |
| `403` | `You are not allowed to chat in this conversation` | before | `conversation_id` belongs to another user. |
| `403` | `User is not a member of the organization` | before | `organization_id` (or the assistant's organization, or the organization of the conversation) is not one of the user's organizations. |
| `403` | `User does not have write permissions in the organization` | before | The user is a guest in the organization. |
| `404` | `Assistant not found` | before | No assistant with `assistant_id`. |
| `404` | `model not found` | before | No model with the id in `model` (or the assistant's model). |
| `422` | list of field errors | before | Invalid body, for example neither `assistant_id` nor `model` given (`Either assistant_id or model is required`). |
| `422` | `max_tokens must be between 1 and the maximum allowed tokens for the model` | before | `max_tokens` out of range. |
| `422` | `the conversation belongs to another organization` | before | The conversation was created in an organization and the request names another one (`organization_id`, or the organization of `assistant_id`). |
| `422` | `invalid base64 data in attachment '<filename>'` | before | The `data` of an attachment is not valid base64. |
| `422` | `invalid data URI in attachment '<filename>'` | before | The `data` of an attachment starts with `data:` but has no `,` before the content. |
| `422` | `file is too large` | before | An attachment exceeds the upload limit. |
| `422` | `storage quota exceeded`, `organization storage quota exceeded` | before | An attachment does not fit into the storage quota. |
| `422` | `Assistant skill '/<slug>' requires unavailable tools: ...` | before | The activated skill needs a tool that is not available. |
| `422` | `model is deprecated` | before | The model (or the assistant's model) is retired. |
| `422` | `model is disabled` | before | The model (or the assistant's model) is switched off by the administrator. |
| `422` | `model is not available` | before | The model (or the assistant's model) is an automatic model and none of the models it chooses from is available. |
| `422` | `model does not have the required capability` | during | The model is not a chat model. |
| `422` | `not enough user credit`, `not enough organization credit` | during | See [Organizations and credits](#organizations-and-credits). |
| `422` | `max iterations reached` | during | The model kept calling tools beyond the allowed number of rounds. |
| `403` | `llm: refusal from provider` | during | The model provider refused to answer. |
| `404` | `organization membership not found` | during | The user left the organization the request runs in. |
| `429` | rate limit message | before | Too many requests, see [Rate limits](conventions.md#rate-limits). |
| `500` | `llm: error on provider side`, `llm: error in chat stream`, `llm: max token reached` | during | The model provider failed, or the answer hit the output token limit. |
| `529` | `llm: provider overload error` | during | The model provider is overloaded; retry later. |
| `500` | `An error occurred during chat processing` | during | Unexpected error (without `stream`). |

If an error interrupts an answer, the part generated so far is kept in the conversation.

### Examples

Ask a model a question and continue the same conversation:

```bash
curl -X POST "https://ayeto.ai/api/v3/chat" \
  -H "uni-api-key: $AYETO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "conversation_id": "3f6c1a9e-8b2d-4e57-9a1c-2d7e5b8f4a10",
    "model": "gpt-5-mini",
    "message": "Summarize the benefits of unit tests in three bullet points."
  }'
```

```json
{
  "id": "b1e4c7d2-5a3f-4c89-8e6b-0f2a9d1c7e35",
  "timestamp": 1759406412345,
  "role": "assistant",
  "content": "- **Catch regressions early** ...\n- **Document behaviour** ...\n- **Make refactoring safe** ...",
  "reasoning_content": "",
  "model": "gpt-5-mini",
  "attachments": [],
  "tool_runs": {}
}
```

```bash
curl -X POST "https://ayeto.ai/api/v3/chat" \
  -H "uni-api-key: $AYETO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "conversation_id": "3f6c1a9e-8b2d-4e57-9a1c-2d7e5b8f4a10",
    "model": "gpt-5-mini",
    "message": "Now give an example for the second point in Python."
  }'
```

Talk to an assistant in an organization and attach a document:

```bash
curl -X POST "https://ayeto.ai/api/v3/chat" \
  -H "uni-api-key: $AYETO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "conversation_id": "8d2f0b6a-1c4e-4f7a-b3d9-6e5a2c1f0b87",
    "assistant_id": "c7a9e2f1-4b6d-4d38-9f0e-1a2b3c4d5e6f",
    "organization_id": "5e8b1d3c-7a2f-4c6e-8d9b-0a1f2e3d4c5b",
    "message": "List the payment terms in this contract.",
    "remove_tool_calls": true,
    "attachments": [
      {
        "filename": "contract.txt",
        "data": "data:text/plain;base64,UGF5bWVudCBkdWUgd2l0aGluIDMwIGRheXMu",
        "mime_type": "text/plain"
      }
    ]
  }'
```

Generate an image in a model conversation:

```bash
curl -X POST "https://ayeto.ai/api/v3/chat" \
  -H "uni-api-key: $AYETO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "conversation_id": "0a7c3e5f-9b1d-4a26-8c4e-7f6d5b3a2c19",
    "model": "gpt-5-mini",
    "dynamic_tools": true,
    "message": "Draw a watercolor lighthouse at sunset."
  }'
```

```json
{
  "id": "e2d4f6a8-0b1c-4d3e-9f5a-7b6c8d9e0f12",
  "timestamp": 1759406530112,
  "role": "assistant",
  "content": "\n\n### calling AI tool ...\n### Generating image\n...\n![image.png](https://ayeto.ai/api/v1/public/file/read?id=9c1e...&secret=4f7a...)\n\n***\nHere is your watercolor lighthouse at sunset.",
  "reasoning_content": "",
  "model": "gpt-5-mini",
  "attachments": [
    {
      "filename": "image.png",
      "content_type": "image/png",
      "file_url": "https://ayeto.ai/api/v1/public/file/read?id=9c1e...&secret=4f7a...",
      "name": null,
      "description": null
    }
  ],
  "tool_runs": {}
}
```

Python (`requests`):

```python
import base64
import os
import uuid

import requests

API = "https://ayeto.ai/api/v3"
HEADERS = {"uni-api-key": os.environ["AYETO_API_KEY"]}

conversation_id = str(uuid.uuid4())  # keep it to continue the conversation


def chat(message, attachments=None):
    response = requests.post(
        f"{API}/chat",
        headers=HEADERS,
        json={
            "conversation_id": conversation_id,
            "model": "gpt-5-mini",
            "message": message,
            "remove_tool_calls": True,
            "attachments": attachments,
        },
        timeout=600,  # answers that use tools can take minutes
    )
    if response.status_code != 200:
        raise RuntimeError(f"{response.status_code}: {response.json()['detail']}")
    return response.json()


with open("invoice.pdf", "rb") as f:
    pdf = f.read()

answer = chat(
    "What is the total amount of this invoice?",
    attachments=[{
        "filename": "invoice.pdf",
        "data": "data:application/pdf;base64," + base64.b64encode(pdf).decode(),
        "mime_type": "application/pdf",
    }],
)
print(answer["content"])
print(chat("And the due date?")["content"])
```

JavaScript (`fetch`):

```javascript
const API = "https://ayeto.ai/api/v3";
const conversationId = crypto.randomUUID(); // keep it to continue the conversation

async function chat(message) {
  const response = await fetch(`${API}/chat`, {
    method: "POST",
    headers: {
      "uni-api-key": process.env.AYETO_API_KEY,
      "Content-Type": "application/json",
    },
    body: JSON.stringify({
      conversation_id: conversationId,
      assistant_id: "c7a9e2f1-4b6d-4d38-9f0e-1a2b3c4d5e6f",
      message,
    }),
  });
  const body = await response.json();
  if (!response.ok) {
    throw new Error(`${response.status}: ${JSON.stringify(body.detail)}`);
  }
  return body;
}

const answer = await chat("Draft a short reply to a customer asking about delivery times.");
console.log(answer.content);
```

To receive the answer while it is generated, set `"stream": true` and read the response
as described in [Streaming](streaming.md#reading-a-stream).
