Text to speech
Stream Arabic speech from text with the audio endpoint.
Send POST /v1/audio/speech with a Bearer key that has the tts scope. The endpoint follows the familiar OpenAI audio-speech request shape and streams the resulting audio in the response body.
curl --fail --show-error https://api.sawtakarabi.ai/v1/audio/speech \
-H "Authorization: Bearer $SAWTAK_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "arabic-tts-1",
"input": "مرحباً بك",
"response_format": "wav"
}' \
--output speech.wav
Request fields
| Field | Required | Description |
|---|---|---|
model | Yes | The Arabic speech model identifier. |
input | Yes | Text to synthesize. |
voice | No | A voice ID returned by the voice catalog. Omit it to use the default voice. |
response_format | No | pcm for raw audio, or wav for a WAV stream. |
sample_rate | No | Requested PCM output rate when supported. |
enhance_pronunciation | No | Optional boolean (true/false). Applies dialect-conditioned automatic diacritization (tashkeel) to Arabic text before synthesis to improve pronunciation accuracy and articulation. Preserves existing marks. Defaults to false. |
Text preparation
The service validates the request before it is sent to a speech worker. Split unusually long passages at natural sentence boundaries, and keep words intact. Text normalization runs automatically by default. Use enhance_pronunciation: true to apply dialect-aware diacritization (tashkeel) for optimal pronunciation.
Use a custom voice
Create a private voice through Voice cloning, wait until its status is ready, then place its ID in voice. A voice that is not available to the authenticated account returns an error instead of silently falling back to another voice.
For response failures and retries, see Errors and Request limits.