Text to speech API: your voices, from your code
A simple REST API: send a text and one of your voices, get the MP3 back. Same credits as the web app, no developer fee and no per-call charge.
- Base URL
- https://voizum.com/api/v1
- Authentication
- Bearer sk_voizum_…
- Output
- MP3
- Languages
- es · en · de · fr · pt · it · ru
- Price
- 60 credits / 1,000 characters (min. 100)
For AI assistants and tools
Three doors, depending on who connects. They all lead to the same place: your voices and your credits.
- Claude, ChatGPT, Cursor… no code
- The MCP connector: add it to the assistant and ask for audio in the chat.
- https://voizum.com/mcp
- ChatGPT Actions, Postman, client generators
- Import the OpenAPI 3.1 spec: every endpoint and field described.
- https://voizum.com/openapi.json
- Agents that read docs
- llms.txt: the whole contract in plain text, made for language models.
- https://voizum.com/llms.txt
Get started in three steps
- Create your account and your key in the API section of the web app. The key starts with sk_voizum_ and is shown once: keep it secret, never in browser code.
- Pick a voice: clone your own or add one from the library to “My voices”. GET /voices gives you its id.
- Create the audio with POST /tts, poll its status every few seconds and download it when it's ready (or get a webhook).
A complete example
Create, wait, download. The API is asynchronous: a long audio takes time and this way nothing has to stay connected.
curl
# 1. Create the audio
curl -X POST https://voizum.com/api/v1/tts \
-H "Authorization: Bearer $VOIZUM_API_KEY" \
-H "Content-Type: application/json" \
-d '{"text": "Hi, this is Voizum.", "voice_id": "VOICE_ID"}'
# → {"job_id": "…", "status": "queued", "credits_charged": 100, "eta_seconds": 20}
# 2. Check the status (until "done")
curl https://voizum.com/api/v1/tts/JOB_ID -H "Authorization: Bearer $VOIZUM_API_KEY"
# 3. Download the MP3
curl -L -o audio.mp3 "https://voizum.com/api/v1/tts/JOB_ID/audio?download=1" -H "Authorization: Bearer $VOIZUM_API_KEY"Python
import os, time, requests
API = "https://voizum.com/api/v1"
H = {"Authorization": f"Bearer {os.environ['VOIZUM_API_KEY']}"}
job = requests.post(f"{API}/tts", headers=H, json={"text": "Hi, this is Voizum.", "voice_id": "VOICE_ID"}).json()
while True:
status = requests.get(f"{API}/tts/{job['job_id']}", headers=H).json()
if status["status"] in ("done", "error", "canceled"):
break
time.sleep(3)
if status["status"] == "done":
mp3 = requests.get(f"{API}/tts/{job['job_id']}/audio?download=1", headers=H)
open("audio.mp3", "wb").write(mp3.content)JavaScript (Node 18+)
const API = "https://voizum.com/api/v1";
const H = { Authorization: `Bearer ${process.env.VOIZUM_API_KEY}`, "Content-Type": "application/json" };
const job = await (await fetch(`${API}/tts`, {
method: "POST", headers: H,
body: JSON.stringify({ text: "Hi, this is Voizum.", voice_id: "VOICE_ID" }),
})).json();
let status;
do {
await new Promise((r) => setTimeout(r, 3000));
status = await (await fetch(`${API}/tts/${job.job_id}`, { headers: H })).json();
} while (!["done", "error", "canceled"].includes(status.status));
if (status.status === "done") {
const mp3 = await fetch(`${API}/tts/${job.job_id}/audio?download=1`, { headers: H });
await (await import("node:fs/promises")).writeFile("audio.mp3", Buffer.from(await mp3.arrayBuffer()));
}Endpoints
| Method and path | What it does |
|---|---|
| GET /status | Service health. No key needed. |
| GET /account | Your credit balance and the per-character price. |
| GET /voices | The voices in your library: their id is the voice_id. |
| POST /tts | Create an audio from a text. |
| GET /tts/{job_id} | Job status: queued, processing, done, error or canceled. |
| GET /tts/{job_id}/audio | The finished MP3 (redirects to the file; ?download=1 to download). |
| DELETE /tts/{job_id} | Cancel a job that hasn't started: its credits are returned. |
| GET /tts | Your latest jobs, paginated. |
| POST /tts/batch | Several separate audios in one request, with a single minimum (2 to 100). |
| POST /tts/dialogue | One audio with several voices taking turns (up to 5 voices). |
What POST /tts accepts
| Field | What it is |
|---|---|
| text | Required. The text, up to 600,000 characters. |
| voice_id | Required. A voice from your library (GET /voices). |
| speed | Optional. Speed, 0.5 to 2 (1 = the voice's natural pace). |
| pause_ms | Optional. Pause between sentences, 0 to 800 ms. |
| language | Optional. auto, es, en, de, fr, pt, it, ru. Without it, detected from the text. |
| loudness_normalization | Optional. Levels the volume so every voice sounds equally loud. |
| webhook_url | Optional. A public https URL we POST to when the audio is finished. |
| Idempotency-Key | Optional header. Retry with the same one and you get the same audio, never charged twice. |
Webhook
With webhook_url there's no need to poll: when it finishes we POST a generation.finished event (3 attempts, 10 s timeout each). A 4xx from your server is not retried.
The event is not signed: treat it as a cue to call GET /tts/{job_id} with your key, which is the source of truth, not as proof that the audio is ready.
{
"event": "generation.finished",
"id": "JOB_ID",
"status": "completed",
"voice_id": "VOICE_ID",
"duration_seconds": 12.4,
"credits": 100,
"error": null
}Limits
20 creations per minute (POST /tts, /tts/batch and /tts/dialogue share it) and 300 requests per minute per key overall. Creation responses have X-RateLimit-Limit and X-RateLimit-Remaining, and every 429 has Retry-After with the seconds to wait.
The MP3 is kept for 4 days; after that GET /tts/{job_id}/audio returns 404, so download and store it yourself.
One audio takes up to 600,000 characters. A batch, 2 to 100 audios and 1,200,000 characters in total. A dialogue, up to 100 turns and 5 different voices.
Errors
Always the same shape: { "error": { "code", "message" } }, sometimes with details (for example, how many credits are missing).
| Code | HTTP | What happened |
|---|---|---|
| no_autorizado | 401 | Missing or invalid API key. |
| parametros_invalidos | 400 | Missing fields or invalid values. |
| cuerpo_invalido | 400 | The body is not valid JSON. |
| texto_demasiado_largo | 400 | The text exceeds the character limit. |
| lote_invalido | 400 | A batch takes 2 to 100 items. |
| dialogo_invalido | 400 | A dialogue takes 2 to 100 turns and up to 5 voices. |
| webhook_invalido | 400 | webhook_url is not https or not a public host. |
| voz_sin_muestra | 400 | That voice can't be used yet. |
| saldo_insuficiente | 402 | Not enough credits (the response says how many are needed). |
| requiere_compra | 403 | Batches need an account that has bought credits. |
| voz_no_encontrada | 404 | The voice is not in your library. |
| no_encontrado | 404 | No job with that id (or it isn't yours). |
| limite_peticiones | 429 | Too many requests: wait for Retry-After. |
| demasiados_en_cola | 429 | Too many jobs in progress: wait for some to finish. |
| mantenimiento | 503 | Maintenance pause: retry later. |
| servicio_no_disponible | 503 | Service temporarily unavailable: retry. |
Questions
- Is there a Python or JavaScript SDK?
- You don't need one: it's a few REST calls with JSON that work with any HTTP client (requests, fetch, curl). If you want a generated client, the OpenAPI spec builds one in your language.
- Is there real-time streaming?
- No. The API is asynchronous: you create the audio, check its status and download the finished MP3. It's built for narration, videos, courses and content, not live conversation.
- Can I use the library voices?
- Yes: add them to “My voices” in the web app first and they'll show up in GET /voices. The API only uses the voices in your library.
- Can I clone a voice through the API?
- No: you clone in the web app (upload the recording once) and then use it through the API with its voice_id.
- How much does it cost?
- Same as the web app: 60 credits per 1,000 characters, with a minimum of 100 per request. No monthly fee and no per-call charge. If an audio fails, its credits are returned.
- What format is the audio?
- MP3.
Type your text and hear it in a real voice
The API spends credits from a pack (60 per 1,000 characters) and needs a Google sign-in. Packs from $6.99, no subscription.