---
title: Text to speech
description: Stream Arabic speech from text with the audio endpoint.
---

Send `POST /v1/audio/speech` with a Bearer key that has the `tts` scope. The endpoint follows the familiar OpenAI audio-speech request shape and streams the resulting audio in the response body.

```bash
curl --fail --show-error https://api.sawtakarabi.ai/v1/audio/speech \
  -H "Authorization: Bearer $SAWTAK_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "arabic-tts-1",
    "input": "مرحباً بك",
    "response_format": "wav"
  }' \
  --output speech.wav
```

## Request fields

| Field | Required | Description |
| --- | --- | --- |
| `model` | Yes | The Arabic speech model identifier. |
| `input` | Yes | Text to synthesize. |
| `voice` | No | A voice ID returned by the [voice catalog](/docs/voices). Omit it to use the default voice. |
| `response_format` | No | `pcm` for raw audio, or `wav` for a WAV stream. |
| `sample_rate` | No | Requested PCM output rate when supported. |
| `enhance_pronunciation` | No | Optional boolean (`true`/`false`). Applies dialect-conditioned automatic diacritization (tashkeel) to Arabic text before synthesis to improve pronunciation accuracy and articulation. Preserves existing marks. Defaults to `false`. |

## Text preparation

The service validates the request before it is sent to a speech worker. Split unusually long passages at natural sentence boundaries, and keep words intact. Text normalization runs automatically by default. Use `enhance_pronunciation: true` to apply dialect-aware diacritization (tashkeel) for optimal pronunciation.

## Use a custom voice

Create a private voice through [Voice cloning](/docs/voice-cloning), wait until its status is ready, then place its ID in `voice`. A voice that is not available to the authenticated account returns an error instead of silently falling back to another voice.

For response failures and retries, see [Errors](/docs/errors) and [Request limits](/docs/limits).
