# Files and media

> Convert documents, spreadsheets, images and audio to text, and turn text into MP3 speech.

Two utility endpoints that work without a conversation: the data loader turns a file
into plain text (the same conversion AYETO uses for knowledge base uploads), and text to
speech turns text into an MP3 file. Both are paid from the user's
[personal credits](account.md#credits), or from the user's credit in an organization
when the request carries an `organization_id`.

## Organizations and credits

Both endpoints accept an optional `organization_id`. The request then runs in that
organization, the same way a [chat](chat.md#organizations-and-credits) request does: the
user's credit in the organization pays for it (the organization's shared balance or the
user's individual budget, see [credits](account.md#credits)) instead of the personal
credits. The key's user must be an admin or a member of the organization; guests and
users outside it get `403`. Without `organization_id` the request is charged to the
personal credits. The ids of the user's organizations are returned by
[list organization memberships](account.md#list-organization-memberships).

## Convert a file to text

| | |
|---|---|
| Endpoint | `POST /api/v3/data-loader/load` |
| Scope | `ayeto.data_loader` |
| Rate limit | [default](conventions.md#rate-limits) |

Send a file as base64; the response contains its text. Depending on the file type the
text is extracted directly or read by an AI model:

| File type | MIME types | How the text is obtained |
|---|---|---|
| Plain text, CSV, Markdown, HTML, source code | any `text/*`, plus `application/json`, `application/ld+json`, `application/xml`, `application/xhtml+xml`, `application/javascript`, `application/ics` | Decoded as is. UTF-8 is detected; other encodings are guessed. |
| PDF | `application/pdf` | The text layer is extracted, prefixed with `PDF page: N` for every page. When a PDF has too little text (a scanned document), its pages are rendered and read by a vision model. |
| Word | `application/vnd.openxmlformats-officedocument.wordprocessingml.document` (`.docx`), `application/msword` (`.doc`) | Text extracted. |
| Excel | `application/vnd.openxmlformats-officedocument.spreadsheetml.sheet` (`.xlsx`), `application/vnd.ms-excel` (`.xls`) | One section per sheet. For `.xlsx`, cell values with their addresses, formulas and named ranges; for `.xls`, cell values only. Very large workbooks are cut after 100,000 non-empty cells, with a note at the end. |
| Images | any `image/*` | Text in the image is read by a vision model. When the image contains (almost) no text, the model describes the image instead. |
| Audio | `audio/mpeg` (`.mp3`), `audio/wav`, `audio/wave`, `audio/x-wav` (`.wav`) | Transcribed by a speech-to-text model. Long recordings are split and transcribed in parts. |

Other file types (archives, video, other audio formats, presentations) are not
converted: the request is refused with `422` `unsupported file type '<mime type>'`
naming the type the file was processed as, and nothing is charged.

The MIME type is taken from, in this order: the `mime_type` field, the header of a data
URL in `data`, and finally detection from the file content. Send the type explicitly
when you know it, detection is not reliable for every format.

The endpoint sets no size limit of its own, but the whole file travels base64-encoded in
one JSON body and is converted while the request is open. Large scanned PDFs, images and
long recordings take time; keep files reasonably small and allow for a long response
time.

### Billing

Direct text extraction (text, PDF with a text layer, Word, Excel) uses no AI model.
Vision (images, scanned PDFs) and speech to text (audio) are charged at the price of the
models used. The server can also be configured with a minimum price per call; when a
call costs less than the minimum, the difference is charged as a fee. The credits a call
consumed are returned in `credits`.

Results of vision reading are cached for 24 hours: converting the same image or scanned
PDF again within that time does not call the model again and costs only the minimum fee,
if one is set.

The check whether the user (or, with `organization_id`, the organization) has credits
happens before the conversion; see [credits](account.md#credits). The call is charged to
the organization when `organization_id` is set, see [organizations](#organizations-and-credits).

### Request

| Field | Type | Required | Description |
|---|---|---|---|
| `data` | string | Yes | The file, base64-encoded. A data URL (`data:application/pdf;base64,JVBERi0x...`) is also accepted. |
| `mime_type` | string | No | MIME type of the file. Overrides the type in a data URL and the detection. |
| `organization_id` | UUID | No | Run the conversion in this organization and charge it to the user's credit there, see [organizations](#organizations-and-credits). The user must be an admin or member (not a guest). Default: none (personal credits). |

### Response

`200 OK` with:

| Field | Type | Description |
|---|---|---|
| `content` | string | The extracted text. |
| `size` | integer | For text input, the size of the input in bytes; for converted files, the length of `content` in characters. |
| `mime_type` | string | The MIME type the file was processed as. For an image whose text was read it is `text/plain`; for an image that was described it stays the image type. |
| `credits` | number | Credits the call consumed, including a minimum fee. |

The response also contains the fields `filename`, `url`, `name`, `meta`, `path`,
`extension`, `is_chunk`, `binary` and `binary_data`. For this endpoint they are always
empty or `false`; ignore them.

### Errors

| Status | `detail` | Cause |
|---|---|---|
| `401` | `API key is invalid` | The key lacks the `ayeto.data_loader` scope, see [authentication](authentication.md#authentication-errors). |
| `403` | `User is not a member of the organization` | `organization_id` is not one of the user's organizations. |
| `403` | `User does not have write permissions in the organization` | The user is a guest in the organization. |
| `422` | `invalid base64 data` | `data` is not valid base64 or a valid data URL. |
| `422` | `invalid encoding` | A text file could not be decoded. |
| `422` | `unsupported file type '<mime type>'` | The file is of a type that is not converted (see the table above), for example `unsupported file type 'application/zip'`. |
| `422` | `not enough user credit` | The user's personal credits are used up. |
| `422` | `not enough organization credit` | With `organization_id`: the user's credit in the organization is used up. |
| `422` | validation error | `data` is missing. |
| `500` | `error reading document` | A scanned PDF could not be rendered. |
| `500` | `Error while processing Word document` | The Word file could not be read. |
| `500` | `Error while processing Excel file` | The Excel file could not be read. |
| `500` | `Error while processing audio file`, `Error while processing WAV audio file` | The audio could not be converted or transcribed. |

Authentication, rate limit and server errors are described in
[conventions](conventions.md#errors).

### Example

```bash
curl -X POST "https://ayeto.ai/api/v3/data-loader/load" \
  -H "uni-api-key: $AYETO_API_KEY" \
  -H "Content-Type: application/json" \
  -d "{\"data\": \"$(base64 -w0 invoice.pdf)\", \"mime_type\": \"application/pdf\"}"
```

```json
{
  "content": "PDF page: 1\nInvoice 2026-0412\nNorthwind Trading s.r.o.\nTotal due: 12 400 CZK\n\n",
  "size": 76,
  "filename": "",
  "url": "",
  "name": "",
  "meta": "",
  "path": "",
  "extension": "",
  "is_chunk": false,
  "mime_type": "application/pdf",
  "binary": false,
  "binary_data": "",
  "credits": 0.0
}
```

## Text to speech

| | |
|---|---|
| Endpoint | `POST /api/v3/tts/mp3` |
| Scope | `ayeto.tts` |
| Rate limit | [default](conventions.md#rate-limits) |

Converts text to speech and returns an MP3 file.

The text is split into sentences (at line breaks and at `. `) and each sentence is
spoken separately; the audio of all sentences is joined into one file. A single
sentence must not be longer than 4,096 characters.

### Billing

Each sentence is charged by its number of characters at the price of the
text-to-speech model (see the `tts` type in [models](models-and-tools.md#model-types)).
A sentence spoken with the same voice within the last hour is reused and not charged
again. Speech is paid from the user's [personal credits](account.md#credits), or from
the user's credit in an organization when `organization_id` is set (see
[organizations](#organizations-and-credits)). The balance is checked before each sentence is
generated.

### Request

| Field | Type | Required | Description |
|---|---|---|---|
| `text` | string | Yes | The text to speak. |
| `voice` | string | No | One of `alloy`, `echo`, `fable`, `onyx`, `nova`, `shimmer`. Default: the server's default voice (`onyx` unless changed by the administrators). |
| `filename` | string | No | File name for the `Content-Disposition` header. Default `audio.mp3`. |
| `organization_id` | UUID | No | Generate the speech in this organization and charge it to the user's credit there, see [organizations](#organizations-and-credits). The user must be an admin or member (not a guest). Default: none (personal credits). |

### Response

`200 OK` with the MP3 file as the body (`Content-Type: audio/mpeg`) and a
`Content-Disposition: attachment` header carrying the file name.

### Errors

| Status | `detail` | Cause |
|---|---|---|
| `401` | `API key is invalid` | The key lacks the `ayeto.tts` scope, see [authentication](authentication.md#authentication-errors). |
| `403` | `permission denied` | The user's account is not allowed to use text to speech. |
| `403` | `User is not a member of the organization` | `organization_id` is not one of the user's organizations. |
| `403` | `User does not have write permissions in the organization` | The user is a guest in the organization. |
| `422` | validation error | `text` is missing, or `voice` is not one of the listed voices. |
| `422` | `not enough user credit` | The user's personal credits are used up. |
| `422` | `not enough organization credit` | With `organization_id`: the user's credit in the organization is used up. |
| `500` | `No audio generated` | `text` is empty. |
| `500` | `Failed to generate audio` | Speech generation failed, for example because a sentence is too long. |

Authentication, rate limit and server errors are described in
[conventions](conventions.md#errors).

### Example

```bash
curl -X POST "https://ayeto.ai/api/v3/tts/mp3" \
  -H "uni-api-key: $AYETO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"text": "Your order has shipped. It will arrive on Friday.", "voice": "nova", "filename": "order-update.mp3"}' \
  --output order-update.mp3
```
