Skip to content
صوتك عربي

Voice cloning

Create and use an authorized cloned voice.

The Voice Cloning API allows you to create custom voice clones from audio samples, monitor their processing status, and use the resulting voice ID for speech synthesis.

Authentication and Scope

All voice cloning endpoints require authentication using a Bearer token in the HTTP Authorization header:

Authorization: Bearer <API_KEY>
  • Required Scope: The API key must include the voices scope. Requests missing this scope receive HTTP 403 Forbidden (insufficient_scope).
  • Anonymous Access: Keyless or anonymous requests are not permitted and return HTTP 401 Unauthorized (auth_missing).

Clone a Voice

To create a new cloned voice, submit a multipart/form-data request to the voices endpoint:

POST /v1/voices
Content-Type: multipart/form-data

Form Fields

FieldTypeRequiredDescription
filesbinaryYesReference audio file. Exactly one audio file must be supplied in this field.
namestringYesDisplay name for the cloned voice.
labelsstring (JSON)YesA JSON-encoded object containing categorization metadata. Must include dialect.
descriptionstringNoOptional text description (maximum 180 characters).
continuation_filebinaryNoOptional secondary audio clip providing delivery context. Must be paired with continuation_text.
continuation_textstringNoSpoken transcript matching continuation_file. Must be paired with continuation_file.

Labels JSON Object

The labels field must be passed as a serialized JSON object with the following properties:

PropertyTypeRequiredDescription
dialectstringYesDialect identifier. Must match a supported dialect code.
genderstringNoVoice gender. Allowed values: male, female, neutral (defaults to male).
agestringNoVoice age group. Allowed values: young, middle, old (defaults to middle).
use_casestringNoIntended application. Allowed values: conversational, narration, advertisement, or empty string "" (defaults to "").

Example Request

curl -X POST "https://api.sawtakarabi.ai/v1/voices" \
  -H "Authorization: Bearer <API_KEY>" \
  -F "name=Custom Voice" \
  -F 'labels={"dialect":"saudi-najdi","gender":"male","age":"middle","use_case":"conversational"}' \
  -F "description=Cloned voice for conversational interactions" \
  -F "files=@reference_audio.wav"

Response

A successful request returns HTTP 201 Created with the assigned voice ID and initial readiness status:

{
  "object": "voice",
  "id": "7b8e62c1-9689-4a7b-a320-912b7a94f6f4",
  "status": "processing"
}

If a continuation sample was provided, the response also includes continuation_status:

{
  "object": "voice",
  "id": "7b8e62c1-9689-4a7b-a320-912b7a94f6f4",
  "status": "processing",
  "continuation_status": "processing"
}

Readiness Lifecycle

Voice cloning is processed asynchronously. After creation, check readiness by fetching the voice resource:

GET /v1/voices/{voice_id}
Authorization: Bearer <API_KEY>

Lifecycle States

The status field indicates whether the voice is ready for speech generation:

StatusDescription
processingThe audio clip is undergoing processing. The voice cannot be used for synthesis yet.
readyProcessing completed successfully. The voice ID is ready for speech synthesis.
failedProcessing could not be completed. The voice cannot be used.

When continuation audio is submitted, continuation_status tracks its state independently (processing, ready, or failed). A voice whose main status is ready can still be used even if continuation processing is pending or failed.

Using the Resulting Voice ID

Once status is ready, pass the voice id in the voice parameter of a speech synthesis request:

curl -X POST "https://api.sawtakarabi.ai/v1/audio/speech" \
  -H "Authorization: Bearer <API_KEY>" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "arabic-tts-1",
    "voice": "<VOICE_ID>",
    "input": "مرحبا بك، هذا اختبار للصوت المستنسخ."
  }' \
  --output output.wav

Premature Requests

If speech synthesis is requested while the voice is still processing, the API rejects the request with HTTP 409 Conflict:

{
  "error": {
    "message": "Voice '<VOICE_ID>' is still processing. Its encode is retried on each request; poll GET /v1/voices/<VOICE_ID> for status.",
    "type": "invalid_request_error",
    "code": "voice_processing"
  }
}

Poll GET /v1/voices/{voice_id} until status becomes ready before invoking speech synthesis.

Continuation Control

For voices created with continuation audio, synthesis requests use the continuation context by default. You can disable this behavior by setting "use_continuation": false in the synthesis request payload.

Recovery States and Resource Management

Error Codes and Recovery

HTTP StatusError CodeCause and Recovery
400 Bad Requestbad_requestInvalid multipart body, missing or multiple audio files in files, empty audio, invalid JSON in labels, missing or unsupported dialect, invalid gender/age/use_case, description exceeding 180 characters, or unsupported audio data. Correct the input fields and resubmit.
400 Bad Requestcontinuation_pairOne of continuation_file or continuation_text was provided without the other. Both must be supplied together.
401 Unauthorizedauth_missing / auth_failedMissing or invalid API key. Ensure a valid key is passed in the Authorization header.
402 Payment Requiredinsufficient_balanceAccount balance is insufficient to complete the voice clone. Add credits to your account before retrying.
403 Forbiddeninsufficient_scopeThe API key lacks the voices scope. Generate or use an API key with the required scope.
404 Not Foundvoice_not_foundThe specified voice ID does not exist or is not owned by the authenticated account.
409 Conflictvoice_limit_reachedThe account has reached its maximum allowed number of voices. Delete an existing voice before cloning another.
409 Conflictvoice_processingSpeech synthesis attempted before processing finished. Poll GET /v1/voices/{voice_id} until status is ready.
413 Payload Too Largeclip_too_largeThe uploaded reference clip exceeds maximum allowable upload size. Reduce the file size and retry.
503 Service Unavailablevoices_unconfigured / storage_unavailableThe voice service or storage backend is temporarily unavailable. Retry the request later.

Updating Voice Metadata

You can update the name, description, or labels of an existing cloned voice:

PATCH /v1/voices/{voice_id}
Authorization: Bearer <API_KEY>
Content-Type: application/json

{
  "name": "Updated Voice Name",
  "description": "Updated voice description",
  "labels": {
    "use_case": "narration"
  }
}

Deleting a Voice

To delete a cloned voice and release capacity within your account limit:

DELETE /v1/voices/{voice_id}
Authorization: Bearer <API_KEY>

A successful deletion returns HTTP 204 No Content.