---
title: Voices
description: Find the right voice and use its ID in a request.
---

The Voices API lets you search, filter, and inspect voices from the catalog, audition preview audio, and retrieve the voice ID required for speech synthesis.

## Authentication and Scopes

Access permissions depend on whether the voice is public or private:

- **Public voices**: Can be listed and previewed without authentication. This allows direct audio playback in web browsers and client applications.
- **Private voices**: Require an API key passed in the standard HTTP `Authorization` header:

```http
Authorization: Bearer <API_KEY>
```

- **Required Scope**: The API key must include the `voices` scope (or have unrestricted access). If the scope is missing, the API returns HTTP `403 Forbidden` (`insufficient_scope`).

## List Voices

Retrieve a paginated list of available voices.

```http
GET /v1/voices
```

### Query Parameters

All query parameters are optional:

| Parameter | Type | Default | Description |
| :--- | :--- | :--- | :--- |
| `sort` | string | `newest` | Sort order for results. Supported values: `newest`, `oldest`, `most_liked`. Any other value returns HTTP `400` (`bad_request`). |
| `limit` | integer | `25` | Number of voices to return per page. Must be an integer between `1` and `100` (inclusive). Values outside this range return HTTP `400` (`bad_request`). |
| `after` | string | `null` | Opaque pagination cursor from a previous response (`next_cursor`). Used to fetch the next page of results. |
| `sharing_status` | string | `null` | Filter by voice visibility: `public` or `private`. If omitted, authenticated requests return the caller's own voices plus all public voices. Unauthenticated requests always return only public voices. |
| `dialect` | string | `null` | Filter by dialect label. This parameter can be repeated multiple times (for example, `?dialect=saudi-najdi&dialect=eg-cairene`) to match any of the specified dialects. |
| `gender` | string | `null` | Filter by gender label (for example, `male`, `female`, `neutral`). |
| `age` | string | `null` | Filter by age label (for example, `young`, `middle`, `old`). |
| `use_case` | string | `null` | Filter by use-case label (for example, `conversational`, `narration`, `advertisement`). |
| `search` | string | `null` | Case-insensitive substring search matching against voice names. |

### Response Structure

The endpoint returns a list envelope object:

| Field | Type | Description |
| :--- | :--- | :--- |
| `object` | string | Always `"list"`. |
| `data` | array | An array of voice objects matching the query. |
| `first_id` | string or null | The `id` of the first voice in `data`, or `null` if the list is empty. |
| `last_id` | string or null | The `id` of the last voice in `data`, or `null` if the list is empty. |
| `has_more` | boolean | Indicates whether additional matching voices exist beyond the current page. |
| `next_cursor` | string (optional) | An opaque cursor string included when `has_more` is `true`. Pass this value to the `after` query parameter to fetch the next page. |

## Voice Object

Each item in `data`, as well as the response from `GET /v1/voices/{voice_id}`, contains the following fields:

| Field | Type | Description |
| :--- | :--- | :--- |
| `object` | string | Always `"voice"`. |
| `id` | string | Unique voice identifier. Use this value as the `voice` parameter in speech synthesis requests. |
| `name` | string | Display name of the voice. |
| `status` | string | Readiness status: `"ready"`, `"processing"`, or `"failed"`. Only voices with status `"ready"` can be used for speech synthesis. |
| `labels` | object | Categorization metadata containing `dialect`, `gender`, `age`, and `use_case` string values. |
| `preview_url` | string | URL pointing to the audio preview endpoint for this voice (`/v1/voices/{id}/preview`). |
| `sharing` | object | Visibility details containing `status` (`"public"` or `"private"`) and `liked_by_count` (integer count of likes). |
| `created_at` | string | Creation timestamp in ISO 8601 format (UTC). |
| `description` | string (optional) | Text description of the voice, present if configured. |
| `preview_text` | string (optional) | The spoken text used in the preview audio sample, present if configured. |
| `continuation_status` | string (optional) | Status of continuation data (`"ready"`, `"processing"`, or `"failed"`), present only on voices created with a continuation sample. |

## Retrieve a Single Voice

Fetch metadata and status for a specific voice by its identifier.

```http
GET /v1/voices/{voice_id}
```

### Access and Visibility

- Returns the voice object if the voice is public or owned by the authenticated caller.
- If the voice does not exist, or is private and belongs to another account, the endpoint returns HTTP `404 Not Found` (`voice_not_found`). Private voices belonging to other accounts are indistinguishable from nonexistent voices.

## Preview Voice Audio

Audition a pre-rendered audio sample of a voice before using it.

```http
GET /v1/voices/{voice_id}/preview
```

### Audio Specifications and Headers

- **Audio format**: Standard WAV (`audio/wav`).
- **Cache-Control**: `public, max-age=300`.
- **Public voices**: Unauthenticated access is allowed so that web players and client applications can stream the sample directly.
- **Private voices**: Requires an authenticated API key with the `voices` scope belonging to the voice owner.
- **Rate limit**: Keyless preview requests are limited to `60` requests per minute per IP address. If this limit is exceeded, the endpoint returns HTTP `429 Too Many Requests` with a `Retry-After: 60` header.

## Choosing a Voice for Synthesis

To synthesize speech using a selected voice:

1. **Find a voice**: Call `GET /v1/voices` with appropriate dialect, gender, or use-case filters.
2. **Verify readiness**: Ensure the voice object has `"status": "ready"`. Voices with status `"processing"` or `"failed"` cannot synthesize speech.
3. **Audition the sample**: Stream audio from `preview_url` or `GET /v1/voices/{voice_id}/preview` to evaluate the voice.
4. **Send the synthesis request**: Pass the voice `id` in the `voice` field of your speech synthesis request:

```bash
curl -X POST "https://api.sawtakarabi.ai/v1/audio/speech" \
  -H "Authorization: Bearer <API_KEY>" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "arabic-tts-1",
    "voice": "<VOICE_ID>",
    "input": "مرحبا بك",
    "response_format": "wav"
  }' \
  --output speech.wav
```

If the `voice` parameter is omitted, the speech synthesis endpoint uses the system default voice.

## Error Reference

The Voices API returns standard error responses:

| HTTP Status | Error Code | Description |
| :--- | :--- | :--- |
| `400 Bad Request` | `bad_request` | Invalid parameter, such as an unrecognized `sort` value, `limit` outside `1`..`100`, an unparseable or mismatched pagination cursor, or an invalid `sharing_status`. |
| `401 Unauthorized` | `auth_missing` / `auth_failed` | Missing or invalid API key when accessing private voice resources. |
| `403 Forbidden` | `insufficient_scope` | The provided API key lacks the required `voices` scope. |
| `404 Not Found` | `voice_not_found` | The requested voice ID does not exist, lacks preview audio, or is private and not owned by the caller. |
| `429 Too Many Requests` | `rate_limit_error` | The client exceeded the preview rate limit of 60 requests per minute per IP address. Check the `Retry-After` header. |
| `503 Service Unavailable` | `voices_unconfigured` | The voice catalog service is currently not configured or unavailable. |
