Voice cloning
Create and use an authorized cloned voice.
The Voice Cloning API allows you to create custom voice clones from audio samples, monitor their processing status, and use the resulting voice ID for speech synthesis.
Authentication and Scope
All voice cloning endpoints require authentication using a Bearer token in the HTTP Authorization header:
Authorization: Bearer <API_KEY>
- Required Scope: The API key must include the
voicesscope. Requests missing this scope receive HTTP403 Forbidden(insufficient_scope). - Anonymous Access: Keyless or anonymous requests are not permitted and return HTTP
401 Unauthorized(auth_missing).
Clone a Voice
To create a new cloned voice, submit a multipart/form-data request to the voices endpoint:
POST /v1/voices
Content-Type: multipart/form-data
Form Fields
| Field | Type | Required | Description |
|---|---|---|---|
files | binary | Yes | Reference audio file. Exactly one audio file must be supplied in this field. |
name | string | Yes | Display name for the cloned voice. |
labels | string (JSON) | Yes | A JSON-encoded object containing categorization metadata. Must include dialect. |
description | string | No | Optional text description (maximum 180 characters). |
continuation_file | binary | No | Optional secondary audio clip providing delivery context. Must be paired with continuation_text. |
continuation_text | string | No | Spoken transcript matching continuation_file. Must be paired with continuation_file. |
Labels JSON Object
The labels field must be passed as a serialized JSON object with the following properties:
| Property | Type | Required | Description |
|---|---|---|---|
dialect | string | Yes | Dialect identifier. Must match a supported dialect code. |
gender | string | No | Voice gender. Allowed values: male, female, neutral (defaults to male). |
age | string | No | Voice age group. Allowed values: young, middle, old (defaults to middle). |
use_case | string | No | Intended application. Allowed values: conversational, narration, advertisement, or empty string "" (defaults to ""). |
Example Request
curl -X POST "https://api.sawtakarabi.ai/v1/voices" \
-H "Authorization: Bearer <API_KEY>" \
-F "name=Custom Voice" \
-F 'labels={"dialect":"saudi-najdi","gender":"male","age":"middle","use_case":"conversational"}' \
-F "description=Cloned voice for conversational interactions" \
-F "files=@reference_audio.wav"
Response
A successful request returns HTTP 201 Created with the assigned voice ID and initial readiness status:
{
"object": "voice",
"id": "7b8e62c1-9689-4a7b-a320-912b7a94f6f4",
"status": "processing"
}
If a continuation sample was provided, the response also includes continuation_status:
{
"object": "voice",
"id": "7b8e62c1-9689-4a7b-a320-912b7a94f6f4",
"status": "processing",
"continuation_status": "processing"
}
Readiness Lifecycle
Voice cloning is processed asynchronously. After creation, check readiness by fetching the voice resource:
GET /v1/voices/{voice_id}
Authorization: Bearer <API_KEY>
Lifecycle States
The status field indicates whether the voice is ready for speech generation:
| Status | Description |
|---|---|
processing | The audio clip is undergoing processing. The voice cannot be used for synthesis yet. |
ready | Processing completed successfully. The voice ID is ready for speech synthesis. |
failed | Processing could not be completed. The voice cannot be used. |
When continuation audio is submitted, continuation_status tracks its state independently (processing, ready, or failed). A voice whose main status is ready can still be used even if continuation processing is pending or failed.
Using the Resulting Voice ID
Once status is ready, pass the voice id in the voice parameter of a speech synthesis request:
curl -X POST "https://api.sawtakarabi.ai/v1/audio/speech" \
-H "Authorization: Bearer <API_KEY>" \
-H "Content-Type: application/json" \
-d '{
"model": "arabic-tts-1",
"voice": "<VOICE_ID>",
"input": "مرحبا بك، هذا اختبار للصوت المستنسخ."
}' \
--output output.wav
Premature Requests
If speech synthesis is requested while the voice is still processing, the API rejects the request with HTTP 409 Conflict:
{
"error": {
"message": "Voice '<VOICE_ID>' is still processing. Its encode is retried on each request; poll GET /v1/voices/<VOICE_ID> for status.",
"type": "invalid_request_error",
"code": "voice_processing"
}
}
Poll GET /v1/voices/{voice_id} until status becomes ready before invoking speech synthesis.
Continuation Control
For voices created with continuation audio, synthesis requests use the continuation context by default. You can disable this behavior by setting "use_continuation": false in the synthesis request payload.
Recovery States and Resource Management
Error Codes and Recovery
| HTTP Status | Error Code | Cause and Recovery |
|---|---|---|
400 Bad Request | bad_request | Invalid multipart body, missing or multiple audio files in files, empty audio, invalid JSON in labels, missing or unsupported dialect, invalid gender/age/use_case, description exceeding 180 characters, or unsupported audio data. Correct the input fields and resubmit. |
400 Bad Request | continuation_pair | One of continuation_file or continuation_text was provided without the other. Both must be supplied together. |
401 Unauthorized | auth_missing / auth_failed | Missing or invalid API key. Ensure a valid key is passed in the Authorization header. |
402 Payment Required | insufficient_balance | Account balance is insufficient to complete the voice clone. Add credits to your account before retrying. |
403 Forbidden | insufficient_scope | The API key lacks the voices scope. Generate or use an API key with the required scope. |
404 Not Found | voice_not_found | The specified voice ID does not exist or is not owned by the authenticated account. |
409 Conflict | voice_limit_reached | The account has reached its maximum allowed number of voices. Delete an existing voice before cloning another. |
409 Conflict | voice_processing | Speech synthesis attempted before processing finished. Poll GET /v1/voices/{voice_id} until status is ready. |
413 Payload Too Large | clip_too_large | The uploaded reference clip exceeds maximum allowable upload size. Reduce the file size and retry. |
503 Service Unavailable | voices_unconfigured / storage_unavailable | The voice service or storage backend is temporarily unavailable. Retry the request later. |
Updating Voice Metadata
You can update the name, description, or labels of an existing cloned voice:
PATCH /v1/voices/{voice_id}
Authorization: Bearer <API_KEY>
Content-Type: application/json
{
"name": "Updated Voice Name",
"description": "Updated voice description",
"labels": {
"use_case": "narration"
}
}
Deleting a Voice
To delete a cloned voice and release capacity within your account limit:
DELETE /v1/voices/{voice_id}
Authorization: Bearer <API_KEY>
A successful deletion returns HTTP 204 No Content.