Voices
Find the right voice and use its ID in a request.
The Voices API lets you search, filter, and inspect voices from the catalog, audition preview audio, and retrieve the voice ID required for speech synthesis.
Authentication and Scopes
Access permissions depend on whether the voice is public or private:
- Public voices: Can be listed and previewed without authentication. This allows direct audio playback in web browsers and client applications.
- Private voices: Require an API key passed in the standard HTTP
Authorizationheader:
Authorization: Bearer <API_KEY>
- Required Scope: The API key must include the
voicesscope (or have unrestricted access). If the scope is missing, the API returns HTTP403 Forbidden(insufficient_scope).
List Voices
Retrieve a paginated list of available voices.
GET /v1/voices
Query Parameters
All query parameters are optional:
| Parameter | Type | Default | Description |
|---|---|---|---|
sort | string | newest | Sort order for results. Supported values: newest, oldest, most_liked. Any other value returns HTTP 400 (bad_request). |
limit | integer | 25 | Number of voices to return per page. Must be an integer between 1 and 100 (inclusive). Values outside this range return HTTP 400 (bad_request). |
after | string | null | Opaque pagination cursor from a previous response (next_cursor). Used to fetch the next page of results. |
sharing_status | string | null | Filter by voice visibility: public or private. If omitted, authenticated requests return the caller's own voices plus all public voices. Unauthenticated requests always return only public voices. |
dialect | string | null | Filter by dialect label. This parameter can be repeated multiple times (for example, ?dialect=saudi-najdi&dialect=eg-cairene) to match any of the specified dialects. |
gender | string | null | Filter by gender label (for example, male, female, neutral). |
age | string | null | Filter by age label (for example, young, middle, old). |
use_case | string | null | Filter by use-case label (for example, conversational, narration, advertisement). |
search | string | null | Case-insensitive substring search matching against voice names. |
Response Structure
The endpoint returns a list envelope object:
| Field | Type | Description |
|---|---|---|
object | string | Always "list". |
data | array | An array of voice objects matching the query. |
first_id | string or null | The id of the first voice in data, or null if the list is empty. |
last_id | string or null | The id of the last voice in data, or null if the list is empty. |
has_more | boolean | Indicates whether additional matching voices exist beyond the current page. |
next_cursor | string (optional) | An opaque cursor string included when has_more is true. Pass this value to the after query parameter to fetch the next page. |
Voice Object
Each item in data, as well as the response from GET /v1/voices/{voice_id}, contains the following fields:
| Field | Type | Description |
|---|---|---|
object | string | Always "voice". |
id | string | Unique voice identifier. Use this value as the voice parameter in speech synthesis requests. |
name | string | Display name of the voice. |
status | string | Readiness status: "ready", "processing", or "failed". Only voices with status "ready" can be used for speech synthesis. |
labels | object | Categorization metadata containing dialect, gender, age, and use_case string values. |
preview_url | string | URL pointing to the audio preview endpoint for this voice (/v1/voices/{id}/preview). |
sharing | object | Visibility details containing status ("public" or "private") and liked_by_count (integer count of likes). |
created_at | string | Creation timestamp in ISO 8601 format (UTC). |
description | string (optional) | Text description of the voice, present if configured. |
preview_text | string (optional) | The spoken text used in the preview audio sample, present if configured. |
continuation_status | string (optional) | Status of continuation data ("ready", "processing", or "failed"), present only on voices created with a continuation sample. |
Retrieve a Single Voice
Fetch metadata and status for a specific voice by its identifier.
GET /v1/voices/{voice_id}
Access and Visibility
- Returns the voice object if the voice is public or owned by the authenticated caller.
- If the voice does not exist, or is private and belongs to another account, the endpoint returns HTTP
404 Not Found(voice_not_found). Private voices belonging to other accounts are indistinguishable from nonexistent voices.
Preview Voice Audio
Audition a pre-rendered audio sample of a voice before using it.
GET /v1/voices/{voice_id}/preview
Audio Specifications and Headers
- Audio format: Standard WAV (
audio/wav). - Cache-Control:
public, max-age=300. - Public voices: Unauthenticated access is allowed so that web players and client applications can stream the sample directly.
- Private voices: Requires an authenticated API key with the
voicesscope belonging to the voice owner. - Rate limit: Keyless preview requests are limited to
60requests per minute per IP address. If this limit is exceeded, the endpoint returns HTTP429 Too Many Requestswith aRetry-After: 60header.
Choosing a Voice for Synthesis
To synthesize speech using a selected voice:
- Find a voice: Call
GET /v1/voiceswith appropriate dialect, gender, or use-case filters. - Verify readiness: Ensure the voice object has
"status": "ready". Voices with status"processing"or"failed"cannot synthesize speech. - Audition the sample: Stream audio from
preview_urlorGET /v1/voices/{voice_id}/previewto evaluate the voice. - Send the synthesis request: Pass the voice
idin thevoicefield of your speech synthesis request:
curl -X POST "https://api.sawtakarabi.ai/v1/audio/speech" \
-H "Authorization: Bearer <API_KEY>" \
-H "Content-Type: application/json" \
-d '{
"model": "arabic-tts-1",
"voice": "<VOICE_ID>",
"input": "مرحبا بك",
"response_format": "wav"
}' \
--output speech.wav
If the voice parameter is omitted, the speech synthesis endpoint uses the system default voice.
Error Reference
The Voices API returns standard error responses:
| HTTP Status | Error Code | Description |
|---|---|---|
400 Bad Request | bad_request | Invalid parameter, such as an unrecognized sort value, limit outside 1..100, an unparseable or mismatched pagination cursor, or an invalid sharing_status. |
401 Unauthorized | auth_missing / auth_failed | Missing or invalid API key when accessing private voice resources. |
403 Forbidden | insufficient_scope | The provided API key lacks the required voices scope. |
404 Not Found | voice_not_found | The requested voice ID does not exist, lacks preview audio, or is private and not owned by the caller. |
429 Too Many Requests | rate_limit_error | The client exceeded the preview rate limit of 60 requests per minute per IP address. Check the Retry-After header. |
503 Service Unavailable | voices_unconfigured | The voice catalog service is currently not configured or unavailable. |