Voizum

Text to speech API: your voices, from your code

A simple REST API: send a text and one of your voices, get the MP3 back. Same credits as the web app, no developer fee and no per-call charge.

Base URL
https://voizum.com/api/v1
Authentication
Bearer sk_voizum_…
Output
MP3
Languages
es · en · de · fr · pt · it · ru
Price
60 credits / 1,000 characters (min. 100)

For AI assistants and tools

Three doors, depending on who connects. They all lead to the same place: your voices and your credits.

Claude, ChatGPT, Cursor… no code
The MCP connector: add it to the assistant and ask for audio in the chat.
https://voizum.com/mcp
ChatGPT Actions, Postman, client generators
Import the OpenAPI 3.1 spec: every endpoint and field described.
https://voizum.com/openapi.json
Agents that read docs
llms.txt: the whole contract in plain text, made for language models.
https://voizum.com/llms.txt

Get started in three steps

  1. Create your account and your key in the API section of the web app. The key starts with sk_voizum_ and is shown once: keep it secret, never in browser code.
  2. Pick a voice: clone your own or add one from the library to “My voices”. GET /voices gives you its id.
  3. Create the audio with POST /tts, poll its status every few seconds and download it when it's ready (or get a webhook).

A complete example

Create, wait, download. The API is asynchronous: a long audio takes time and this way nothing has to stay connected.

curl

# 1. Create the audio
curl -X POST https://voizum.com/api/v1/tts \
  -H "Authorization: Bearer $VOIZUM_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"text": "Hi, this is Voizum.", "voice_id": "VOICE_ID"}'
# → {"job_id": "…", "status": "queued", "credits_charged": 100, "eta_seconds": 20}

# 2. Check the status (until "done")
curl https://voizum.com/api/v1/tts/JOB_ID -H "Authorization: Bearer $VOIZUM_API_KEY"

# 3. Download the MP3
curl -L -o audio.mp3 "https://voizum.com/api/v1/tts/JOB_ID/audio?download=1" -H "Authorization: Bearer $VOIZUM_API_KEY"

Python

import os, time, requests

API = "https://voizum.com/api/v1"
H = {"Authorization": f"Bearer {os.environ['VOIZUM_API_KEY']}"}

job = requests.post(f"{API}/tts", headers=H, json={"text": "Hi, this is Voizum.", "voice_id": "VOICE_ID"}).json()
while True:
    status = requests.get(f"{API}/tts/{job['job_id']}", headers=H).json()
    if status["status"] in ("done", "error", "canceled"):
        break
    time.sleep(3)

if status["status"] == "done":
    mp3 = requests.get(f"{API}/tts/{job['job_id']}/audio?download=1", headers=H)
    open("audio.mp3", "wb").write(mp3.content)

JavaScript (Node 18+)

const API = "https://voizum.com/api/v1";
const H = { Authorization: `Bearer ${process.env.VOIZUM_API_KEY}`, "Content-Type": "application/json" };

const job = await (await fetch(`${API}/tts`, {
  method: "POST", headers: H,
  body: JSON.stringify({ text: "Hi, this is Voizum.", voice_id: "VOICE_ID" }),
})).json();

let status;
do {
  await new Promise((r) => setTimeout(r, 3000));
  status = await (await fetch(`${API}/tts/${job.job_id}`, { headers: H })).json();
} while (!["done", "error", "canceled"].includes(status.status));

if (status.status === "done") {
  const mp3 = await fetch(`${API}/tts/${job.job_id}/audio?download=1`, { headers: H });
  await (await import("node:fs/promises")).writeFile("audio.mp3", Buffer.from(await mp3.arrayBuffer()));
}

Endpoints

Method and pathWhat it does
GET /statusService health. No key needed.
GET /accountYour credit balance and the per-character price.
GET /voicesThe voices in your library: their id is the voice_id.
POST /ttsCreate an audio from a text.
GET /tts/{job_id}Job status: queued, processing, done, error or canceled.
GET /tts/{job_id}/audioThe finished MP3 (redirects to the file; ?download=1 to download).
DELETE /tts/{job_id}Cancel a job that hasn't started: its credits are returned.
GET /ttsYour latest jobs, paginated.
POST /tts/batchSeveral separate audios in one request, with a single minimum (2 to 100).
POST /tts/dialogueOne audio with several voices taking turns (up to 5 voices).

What POST /tts accepts

FieldWhat it is
textRequired. The text, up to 600,000 characters.
voice_idRequired. A voice from your library (GET /voices).
speedOptional. Speed, 0.5 to 2 (1 = the voice's natural pace).
pause_msOptional. Pause between sentences, 0 to 800 ms.
languageOptional. auto, es, en, de, fr, pt, it, ru. Without it, detected from the text.
loudness_normalizationOptional. Levels the volume so every voice sounds equally loud.
webhook_urlOptional. A public https URL we POST to when the audio is finished.
Idempotency-KeyOptional header. Retry with the same one and you get the same audio, never charged twice.

Webhook

With webhook_url there's no need to poll: when it finishes we POST a generation.finished event (3 attempts, 10 s timeout each). A 4xx from your server is not retried.

The event is not signed: treat it as a cue to call GET /tts/{job_id} with your key, which is the source of truth, not as proof that the audio is ready.

{
  "event": "generation.finished",
  "id": "JOB_ID",
  "status": "completed",
  "voice_id": "VOICE_ID",
  "duration_seconds": 12.4,
  "credits": 100,
  "error": null
}

Limits

20 creations per minute (POST /tts, /tts/batch and /tts/dialogue share it) and 300 requests per minute per key overall. Creation responses have X-RateLimit-Limit and X-RateLimit-Remaining, and every 429 has Retry-After with the seconds to wait.

The MP3 is kept for 4 days; after that GET /tts/{job_id}/audio returns 404, so download and store it yourself.

One audio takes up to 600,000 characters. A batch, 2 to 100 audios and 1,200,000 characters in total. A dialogue, up to 100 turns and 5 different voices.

Errors

Always the same shape: { "error": { "code", "message" } }, sometimes with details (for example, how many credits are missing).

CodeHTTPWhat happened
no_autorizado401Missing or invalid API key.
parametros_invalidos400Missing fields or invalid values.
cuerpo_invalido400The body is not valid JSON.
texto_demasiado_largo400The text exceeds the character limit.
lote_invalido400A batch takes 2 to 100 items.
dialogo_invalido400A dialogue takes 2 to 100 turns and up to 5 voices.
webhook_invalido400webhook_url is not https or not a public host.
voz_sin_muestra400That voice can't be used yet.
saldo_insuficiente402Not enough credits (the response says how many are needed).
requiere_compra403Batches need an account that has bought credits.
voz_no_encontrada404The voice is not in your library.
no_encontrado404No job with that id (or it isn't yours).
limite_peticiones429Too many requests: wait for Retry-After.
demasiados_en_cola429Too many jobs in progress: wait for some to finish.
mantenimiento503Maintenance pause: retry later.
servicio_no_disponible503Service temporarily unavailable: retry.

Questions

Is there a Python or JavaScript SDK?
You don't need one: it's a few REST calls with JSON that work with any HTTP client (requests, fetch, curl). If you want a generated client, the OpenAPI spec builds one in your language.
Is there real-time streaming?
No. The API is asynchronous: you create the audio, check its status and download the finished MP3. It's built for narration, videos, courses and content, not live conversation.
Can I use the library voices?
Yes: add them to “My voices” in the web app first and they'll show up in GET /voices. The API only uses the voices in your library.
Can I clone a voice through the API?
No: you clone in the web app (upload the recording once) and then use it through the API with its voice_id.
How much does it cost?
Same as the web app: 60 credits per 1,000 characters, with a minimum of 100 per request. No monthly fee and no per-call charge. If an audio fails, its credits are returned.
What format is the audio?
MP3.

Type your text and hear it in a real voice

The API spends credits from a pack (60 per 1,000 characters) and needs a Google sign-in. Packs from $6.99, no subscription.