API reference
Base URL /v1 · JSON in, audio out · $ per million input characters.
Quickstart
- Create an account — you get free characters.
- Create a key on the API keys page.
- Make a request:
curl "{API}/audio/speech" \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "{MODEL}", "voice": "alba", "input": "Hello! Your order ships today."}' \
--output hello.mp3
Authentication
Send your key as a bearer token: Authorization: Bearer sk_live_.... X-API-Key is also accepted. Keys are secret; call the API from your server, never from a browser or mobile app.
sk_test_ keys return real audio for free, limited to 100 characters per request and 100 requests per day.
Using the OpenAI SDKs
The speech endpoint is wire-compatible with OpenAI's. Point the SDK's base URL at us; tts-1, tts-1-hd and gpt-4o-mini-tts are accepted as model names, and OpenAI voice names (alloy, nova, onyx, ...) map to similar stock voices. instructions is accepted and ignored (reported in X-Ignored-Parameters).
from openai import OpenAI
client = OpenAI(base_url="{API}", api_key="sk_live_...")
client.audio.speech.create(model="{MODEL}", voice="alba", input="Hi!").write_to_file("hi.mp3")
Create speech
POST/v1/audio/speech
| Parameter | Type | Default | Description |
|---|---|---|---|
input | string | string[] | required | Text to speak. Up to 10,000 characters (20,000 with stream). Pass an array to keep your own chunk boundaries (e.g. sentences from an LLM). |
voice | string | alba | Stock voice id, your clone id (cloned_...) or an OpenAI voice name. |
model | string | {MODEL} | Optional. OpenAI model names are accepted as aliases. |
response_format | string | mp3 | mp3, opus (Ogg), aac, flac, wav, pcm (24 kHz, 16-bit, mono, little-endian). |
speed | number | 1.0 | 0.5 to 2.0. Pitch is preserved. |
normalize | boolean | true | Turn text into speakable words (see normalization). Disable if you pre-process text yourself. |
pronunciations | array | - | Up to 100 {"word","respelling"} overrides, applied before the built-in dictionary. |
stream | boolean | false | Stream audio as it is generated. mp3, opus, aac, pcm. stream_format: "audio" is an alias. |
delivery | string | bytes | bytes: audio body. url: JSON with a download link valid for 24 hours. json: base64 audio plus usage. |
include_cost | boolean | false | Adds X-Cost-USD and X-Balance-USD headers. JSON deliveries always include a usage object. |
return_normalized_text | boolean | false | Echo the exact text spoken, per chunk (json/url delivery). |
seed | integer | random | Repeatable output for the same input and settings. |
chunk_silence_ms | integer | 80 | Pause inserted between chunks, 0-2000. |
Every response carries X-Request-Id, X-Billable-Characters and X-Voice-Id; non-streamed responses also include X-Audio-Duration-Ms.
Streaming
With "stream": true the first bytes arrive after the first sentence is synthesized, so playback can start long before the full clip is ready. Use pcm for the lowest latency in voice agents (play samples directly), or mp3 for broad player support.
curl -N "{API}/audio/speech" -H "Authorization: Bearer $API_KEY" -H "Content-Type: application/json" \
-d '{"input": "A long story...", "voice": "george", "response_format": "pcm", "stream": true}' \
| ffplay -f s16le -ar 24000 -ac 1 -nodisp -autoexit -
Delivery options
Audio is never stored by default. With "delivery": "url" we keep the file for 24 hours and return:
{
"object": "audio.speech",
"id": "req_...",
"voice": "alba",
"format": "mp3",
"duration_ms": 5120,
"audio": { "url": "{API}/files/aud_...mp3?exp=...&sig=...", "expires_at": 1791763200, "bytes": 41210 },
"usage": { "billable_characters": 84, "normalized_characters": 97, "cost_usd": 0.00042, "balance_usd": 12.34558 }
}
Anyone with the link can download the file until it expires, so treat links like secrets. Not available with zero data retention.
Text normalization
On by default. Numbers, currency ($1,299.99), dates (3/14), times (9:30pm), ordinals, temperatures, phone-style digit runs, acronyms, Roman numerals, markdown, URLs and emoji are converted into words the voice reads naturally. Preview the result for free:
POST/v1/normalize
{ "input": "Dr. Smith paid $42.50 on 3/14.", "pronunciations": [{ "word": "Smith", "respelling": "Smyth" }] }
-> { "normalized": "Doctor Smyth paid forty two dollars and fifty cents on March fourteenth.", "billable_characters": 30, "estimated_cost_usd": 0.00015 }
List voices
GET/v1/voices GET/v1/voices/{id}
Returns your clones first, then the stock catalogue. Filter with ?type=clone or ?type=stock. Every stock voice has a preview_url you can play without authentication.
Voice cloning
POST/v1/voices PATCH/v1/voices/{id} DELETE/v1/voices/{id}
Accept the cloning terms once in the dashboard, then upload 3-60 seconds of clean speech (10-20 s is ideal). Accepted formats: wav, mp3, m4a/aac, ogg/opus, webm, flac, up to 12 MB. Up to 20 voices per account; creating them is free and they cost the same to use as stock voices.
curl "{API}/voices" -H "Authorization: Bearer $API_KEY" \
-F name="Sam narrator" -F [email protected] -F consent=true
-> { "id": "cloned_3f9c...", "object": "voice", "name": "Sam narrator", "type": "clone", ... }
consent=true is your attestation, for this specific recording, that it is your voice or that you have the speaker's documented permission. Your recording is decoded in memory and discarded; we keep only an encrypted voice profile.
Balance & usage
GET/v1/balance
{ "object": "balance", "balance_usd": 12.34, "characters_remaining": 2468000,
"auto_topup": { "enabled": true, "threshold_usd": 5, "amount_usd": 25, "pending": false, "last_failure": null } }
GET/v1/usage?start=2026-10-01&end=2026-10-31&group_by=day — group_by is day, key or none. Totals include requests, errors, characters, audio seconds and cost.
Warmup & cold starts
Most requests are served by a warm instance. A cold instance takes up to ~11 seconds to start. If a request must be fast (a live call is about to begin, say), warm up a few seconds beforehand. It is free and limited to once per minute per key.
POST/v1/warmup?wait=true
{ "object": "warmup", "status": "warm", "waited_ms": 8420, "instance_uptime_s": 9.1 }
Without wait it returns immediately with "warming" or "warm". Instances stay warm for several minutes after their last request. Under bursty load a new request may still land on a fresh instance.
Zero data retention
Turn on ZDR for your whole account (Profile) or per key. With ZDR we write no request logs and store none of your text. Character counts and charges are still recorded because they are your billing record, and /v1/usage keeps working. delivery: "url" is rejected with zdr_conflict.
Errors
Errors use OpenAI's shape and always include a stable code and the request_id:
{ "error": { "message": "This request costs $0.000420 but your balance is $0.000100. Add credits or enable auto top-up.",
"type": "billing_error", "code": "insufficient_balance", "param": null, "request_id": "req_8fK2...", "doc_url": "..." } }
If a request fails on our side (5xx), you are not charged. 429 and 503 responses include Retry-After.
| HTTP | code | What to do |
|---|---|---|
| 400 | invalid_request, invalid_json | Fix the parameter named in param. |
| 400 | empty_input | The input had nothing speakable (only emoji or symbols?). |
| 400 | unsupported_format | Use a listed format; streaming needs mp3, opus, aac or pcm. |
| 400 | model_not_found | Use {MODEL} or an OpenAI TTS model name. |
| 400 | invalid_pronunciation | Words and respellings may contain letters, spaces, apostrophes and hyphens only. |
| 400 | zdr_conflict | Use bytes or json delivery with ZDR. |
| 400 | consent_required | Send consent=true when cloning. |
| 400 | test_key_limit | Test keys allow 100 characters per request. |
| 401 | missing_api_key, invalid_api_key, api_key_revoked | Check the key or create a new one. |
| 402 | insufficient_balance | Add credits or turn on auto top-up. |
| 402 | monthly_cap_reached | Raise or remove the key's monthly cap. |
| 403 | account_suspended | Contact support. |
| 403 | clone_consent_required | Accept the cloning terms in the dashboard. |
| 403 | test_key_not_allowed | Use a live key for cloning. |
| 404 | voice_not_found | List valid ids with GET /v1/voices. |
| 409 | idempotency_conflict, idempotency_in_progress, idempotency_replay | See retries. |
| 409 | voice_limit_reached, key_limit_reached | Delete an existing voice or key. |
| 410 | link_expired | Download links last 24 hours; generate again. |
| 413 | input_too_long, audio_file_too_large | Split the text, or use stream for up to 20,000 characters. |
| 415 | unsupported_audio_type, unsupported_media_type | Upload a supported audio format; send JSON with the right Content-Type. |
| 422 | clone_audio_too_short, clone_audio_too_long, clone_audio_unusable | Upload 3-60 s of clear speech. |
| 429 | rate_limited, concurrency_limited, test_quota_exhausted | Wait for Retry-After seconds. |
| 500 | internal_error | Retry; quote the request id if it persists. Not billed. |
| 502 | engine_error | Retry. Not billed. |
| 503 | engine_unavailable, engine_cold_start_timeout | Retry after Retry-After, or warm up first. Not billed. |
| 504 | upstream_timeout | Retry with shorter input. Not billed. |
Retries & idempotency
Send an Idempotency-Key header (any unique string, e.g. a UUID) to make retries safe for 24 hours:
- Same key, same body, first attempt failed → the retry runs normally.
- Same key while the first is still running →
409 idempotency_in_progress. - Same key after success →
409 idempotency_replay; you are not charged again. Fordelivery: "url"the original result is returned inerror.previous_result. - Same key with a different body →
409 idempotency_conflict.
Retry 429, 500, 502, 503 and 504 with exponential backoff. Don't retry other 4xx errors without changing the request.
Limits
| Characters per request | 10,000 (20,000 streaming) |
| Requests per minute | 120 per key (configurable up to 600) |
| Concurrent requests | 10 per account |
| Cloned voices | 20 per account |
| Clone reference audio | 3-60 s, 12 MB |
| Download links | 24 hours |
| Request logs | 30 days (none with ZDR) |
Need more? Contact us.
OpenAPI spec
The machine-readable spec is at /v1/openapi.json. Import it into Postman, Insomnia or an SDK generator.