Skip to content
صوتك عربي

Quickstart

Make your first Arabic speech request.

Generate Arabic speech from text using the /v1/audio/speech endpoint. This guide walks through authentication, selecting a voice, sending a synthesis request, saving the audio, and handling supported errors.

Set these variables once before running the examples:

export SAWTAK_API_BASE_URL="https://api.sawtakarabi.ai/v1"
export SAWTAK_API_KEY="<API_KEY>"

1. Create an API Key

Requests require a Bearer token passed in the Authorization header.

  • For speech synthesis, your API key must include the tts scope.
  • To query the voice catalog, your API key must include the voices scope.

Pass the key using the SAWTAK_API_KEY environment variable:

-H "Authorization: Bearer $SAWTAK_API_KEY"

2. Choose a Voice

Query the voice catalog using GET $SAWTAK_API_BASE_URL/voices:

curl --fail --show-error "$SAWTAK_API_BASE_URL/voices" \
  -H "Authorization: Bearer $SAWTAK_API_KEY"

The response returns a list envelope (object: "list") containing voice objects in data. Each voice item includes an id, name, status, and labels (such as dialect and gender).

To use a specific voice, pass its id in the voice field of your synthesis request. If you omit the voice field, the endpoint uses the default voice.

3. Issue a Speech Request

Send a POST request to $SAWTAK_API_BASE_URL/audio/speech with the model, text input, and desired response format:

curl --fail --show-error "$SAWTAK_API_BASE_URL/audio/speech" \
  -H "Authorization: Bearer $SAWTAK_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "arabic-tts-1",
    "input": "مرحبا بك",
    "response_format": "wav"
  }' \
  --output speech.wav

To target a specific voice from the catalog, add the voice property with that voice's id.

4. Play or Save the Audio

The speech endpoint streams audio bytes directly into the file specified by --output:

  • WAV (response_format: "wav"): The gateway prepends a standard RIFF/WAV header (audio/wav) to the stream. You can play it immediately with command-line tools:
    afplay speech.wav
    or on Linux:
    aplay speech.wav
  • PCM (response_format: "pcm"): The gateway streams raw 16-bit PCM audio (audio/pcm). Save it to speech.pcm when integrating with streaming decoders or pipelines that consume headerless PCM.

5. Supported Failure Responses

Error responses follow the OpenAI-compatible error envelope containing message, type, code, param, and request_id.

Unknown Voice (400 Bad Request)

Supplying a voice ID that does not exist returns an invalid_request_error with the bad_voice code:

{
  "error": {
    "message": "Unknown voice 'unknown-voice-id'. `voice` must be one of your voice ids from /v1/voices; omit it for the default voice.",
    "type": "invalid_request_error",
    "code": "bad_voice",
    "param": null,
    "request_id": null
  }
}

Insufficient Scope (403 Forbidden)

Using an API key that lacks the required tts scope returns a permission_error:

{
  "error": {
    "message": "Insufficient scope",
    "type": "permission_error",
    "code": "insufficient_scope",
    "param": null,
    "request_id": null
  }
}

Authentication Failure (401 Unauthorized)

Omitting the authorization token or supplying an invalid key returns an authentication_error:

{
  "error": {
    "message": "Authentication failed",
    "type": "authentication_error",
    "code": "auth_failed",
    "param": null,
    "request_id": null
  }
}