Guides
Text to speech API: guide with Python and JS examples
The Voizum team · Updated · 10 min read
A text to speech API turns text into an audio file from your own code or automation, with no one clicking buttons. This guide covers how to pick one for your use case, then walks through a complete, working integration: create the job, check its status, download the MP3. It also covers long scripts, batches, using your own cloned voice and plugging it into n8n, Make or Zapier.
What a text to speech API does and when you need one
A TTS API is an HTTP endpoint: you send text and a voice, and you get audio back. It makes sense as soon as audio becomes part of a process rather than a one-off task. If you narrate one video a month, a web editor is faster. If you narrate one a day, or your product needs to speak, an API saves the copy-paste-download loop every time.
Where we see it used most:
- YouTube and faceless channels: the script comes out of a writing step (often an AI model) and goes straight to narration, then to the video editor.
- E-learning and courses: hundreds of short lessons regenerated whenever the text changes.
- Apps and products: notifications, onboarding, accessibility, read-aloud buttons for articles.
- Batches: product descriptions, audio versions of a blog, phone menus, multilingual versions of the same script.
- No-code automations: n8n, Make or Zapier flows that turn a new row in a spreadsheet into an MP3.
How to choose a text to speech API
Start from what you are building, not from the voice demo. The question that splits the market is latency versus length: a voice assistant needs the first syllable in a few hundred milliseconds and streams audio as it is generated, while a narrator needs a 40-minute script to come out as one clean file. Few APIs are good at both, and the cheapest per character are often the least natural.
| Criterion | Why it matters | What to check |
|---|---|---|
| Latency vs. long audio | Real-time agents need streaming; narration needs long files without seams | Is there streaming? What is the max text per request? |
| Languages | Each language has its own set of voices | Does the same voice speak every language you need? |
| Custom and cloned voices | A recognizable voice is a brand asset | Can you use your own cloned voice through the API? Is there a cap on how many? |
| Price per character | It is how almost every provider ends up billing | Price per 1,000 characters, minimum per request, whether the API costs extra |
| Unused credits | Monthly plans reset; you pay for what you did not use | Do credits roll over, expire or vanish when you cancel? |
| Limits | Requests per minute and jobs in progress shape your architecture | Rate-limit headers, Retry-After, max concurrent jobs |
| Output and storage | You will store, edit or stream the file | Format (MP3, WAV, PCM), how long the provider keeps it |
Get your key and your voice_id
Everything is plain REST with JSON, base URL https://voizum.com/api/v1, and authentication with a Bearer key that starts with sk_voizum_. The flow is always the same: create a job, check its status, download the MP3.
- 1Sign in to Voizum with Google and buy any credit pack: the API key unlocks with both. The API spends purchased credits, never the free monthly ones.
- 2Open the API page and generate your key. It is shown once, so store it right away in an environment variable (VOIZUM_API_KEY) or your secret manager.
- 3Add a voice to your library: clone your own or save one from the public library with Add. The API only works with voices in your library.
- 4List them with GET /voices. Each item has id, name and language; the id is the voice_id you send when generating. You can also copy it in the web: Voices, the ⋯ menu, API ID.
- 5Optional: before a big run, GET /status (public, no key) tells you the service is up, and GET /account shows your credit balance and limits.
Code examples: curl, Python and JavaScript
The examples below work as is once you replace YOUR_VOICE_ID and JOB_ID. The audio URL returned by the status call needs the same Authorization header and redirects to the file, so let your HTTP client follow redirects.
curl
Create the job (202 with job_id and eta_seconds), poll it, then download:
# 1. Create the job
curl -X POST https://voizum.com/api/v1/tts \
-H "Authorization: Bearer $VOIZUM_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: video-042-intro" \
-d '{"text": "Welcome back to the channel.", "voice_id": "YOUR_VOICE_ID"}'
# 2. Check its status (repeat until "done")
curl https://voizum.com/api/v1/tts/JOB_ID -H "Authorization: Bearer $VOIZUM_API_KEY"
# 3. Download the MP3
curl -L -o narration.mp3 https://voizum.com/api/v1/tts/JOB_ID/audio -H "Authorization: Bearer $VOIZUM_API_KEY"Python (requests)
The same flow, waiting while the job is queued or processing:
import os, time, requests
API = "https://voizum.com/api/v1"
H = {"Authorization": "Bearer " + os.environ["VOIZUM_API_KEY"]}
# 1. Create the job
with open("script.txt", encoding="utf-8") as f:
texto = f.read()
r = requests.post(API + "/tts", headers=H, json={"text": texto, "voice_id": "YOUR_VOICE_ID"})
r.raise_for_status()
job = r.json()["job_id"]
# 2. Wait while it is queued or processing
while True:
s = requests.get(API + "/tts/" + job, headers=H).json()
if s["status"] not in ("queued", "processing"):
break
time.sleep(5)
if s["status"] != "done":
raise RuntimeError(s.get("error", s["status"]))
# 3. Download the MP3
audio = requests.get(s["audio_url"], headers=H)
audio.raise_for_status()
with open("narration.mp3", "wb") as f:
f.write(audio.content)JavaScript (Node 18+, ES modules)
Native fetch, no SDK needed:
import fs from "node:fs";
const API = "https://voizum.com/api/v1";
const H = { Authorization: "Bearer " + process.env.VOIZUM_API_KEY, "Content-Type": "application/json" };
// 1. Create the job
const crear = await fetch(API + "/tts", {
method: "POST",
headers: H,
body: JSON.stringify({ text: "Welcome back to the channel.", voice_id: "YOUR_VOICE_ID" }),
});
if (!crear.ok) throw new Error(await crear.text());
const { job_id } = await crear.json();
// 2. Wait while it is queued or processing
let s;
do {
await new Promise((r) => setTimeout(r, 5000));
s = await (await fetch(API + "/tts/" + job_id, { headers: H })).json();
} while (s.status === "queued" || s.status === "processing");
if (s.status !== "done") throw new Error(s.error ?? s.status);
// 3. Download the MP3
const audio = await fetch(s.audio_url, { headers: H });
fs.writeFileSync("narration.mp3", Buffer.from(await audio.arrayBuffer()));Optional parameters
speed (0.5 to 2, default 1), pause_ms (pause between sentences, 0 to 800 ms), language ("auto" or es, en, de, fr, pt, it, ru; by default it is detected from the text) and loudness_normalization (true to make every voice come out at the same level). Status values are queued, processing, done, error and canceled. The full reference is at https://voizum.com/docs, with an OpenAPI spec at https://voizum.com/openapi.json and a plain-text version for AI coding assistants at https://voizum.com/llms.txt.
Long scripts, batches and dialogue
One request takes up to 600,000 characters, roughly 10 hours of speech, and comes back as a single MP3. You do not need to split a long script into chunks and stitch them yourself: the job does that internally and the status call shows progress (parts_done, parts_total, percent).
For many short pieces, use POST /tts/batch: 100 items maximum, 1,200,000 characters in total (about 20 hours). Each item gets its own job_id and its own file, but the 100-credit minimum is charged once for the whole batch instead of once per request. Eight short app prompts cost 800 credits sent one by one and 100 in a batch.
For a conversation in one file (a podcast intro, an ad with two voices) use POST /tts/dialogue with an ordered turns array, up to 5 different voices. It is billed like a normal audio, by the characters of the whole script.
Using your cloned voice through the API
Cloning happens once in the web and costs 199 credits; after that the voice is just another voice_id and every generation is billed like any other voice. There is no clone endpoint in the API on purpose: the recording is the step where you want to listen before committing. Record 15 to 60 seconds of only your voice, cut on a pause rather than mid-word, and talk the way you want the narration to sound, because the intonation is copied too.
Your cloned voices can speak all seven languages, so one English recording also narrates the Spanish, German or French version of your channel. Clone only your own voice or one you have permission to use.
Connecting it to n8n, Make or Zapier
No automation platform needs a special Voizum module: all three have a generic HTTP step, and the API is three plain calls. The only design choice is how you wait for the audio: poll in a loop, or let the API call you back with webhook_url.
The webhook option (POST /tts only) sends a POST with event, id, status (completed or error), voice_id, duration_seconds and credits when the job ends, retried up to three times. It must be a public https URL, so a self-hosted n8n on localhost cannot receive it: poll instead.
n8n
HTTP Request node with Header Auth credentials (Authorization, value Bearer sk_voizum_…), then a loop for the wait:
- HTTP Request: POST https://voizum.com/api/v1/tts with a JSON body containing text and voice_id.
- Wait node for a few seconds, then HTTP Request: GET /tts/{{ $json.job_id }}.
- If node: status is done, continue; queued or processing, back to Wait; error, stop the run.
- HTTP Request on audio_url with the same credentials and the response format set to File; the binary goes on to Drive, S3 or your video step.
Make and Zapier
In Make, the HTTP app's Make a request module sends the POST with the Authorization header. In Zapier, Webhooks by Zapier does the same with a POST or Custom Request action. In both, the cleanest wait is a second scenario or Zap that starts from a custom webhook (Make) or a Catch Hook trigger (Zapier): pass that URL as webhook_url, and when it fires, fetch GET /tts/{id} and download audio_url with your key.
Errors and good practices in production
Every error has the same shape, { error: { code, message } }. Program against code, which is stable, never against message, which is human text and can change.
- Send an Idempotency-Key on every POST /tts (your video or row id works). If a network error makes you retry, you get the same job back with a 200 and are not charged twice.
- Save the job_id as soon as you get it. If your process crashes, GET /tts lists your recent jobs with a cursor, and you can filter with ?status=done.
- Poll every few seconds, or sleep for eta_seconds, rather than in a tight loop.
- Download and keep your own copy: the MP3 is kept for 4 days, then the file is deleted.
- Made a mistake? DELETE /tts/{id} cancels a job that is still queued and refunds it in full. Once it is processing it cannot be cancelled. If a job ends in error, its credits are returned automatically.
- Never call the API from the browser or a mobile app: the key would be visible to anyone. Call it from your server, and have your frontend talk to your server.
- Rotate the key if it may have leaked. Generating a new one revokes the old one immediately, so update the environment variable straight after.
| code | HTTP | What to do |
|---|---|---|
| no_autorizado | 401 | Missing, wrong or revoked key |
| saldo_insuficiente | 402 | Not enough purchased credits; the response says how many you need |
| voz_no_encontrada | 404 | That voice_id is not in your library: add it first |
| texto_demasiado_largo | 400 | Over 600,000 characters: split it |
| limite_peticiones | 429 | Wait the seconds in Retry-After, then retry |
| demasiados_en_cola | 429 | Too many jobs in progress: let some finish |
| servicio_no_disponible | 503 | Retry later with backoff |
How much does a text to speech API cost?
The API spends the same credits as the web: 60 credits per 1,000 characters (about 60 per minute of audio), with a minimum of 100 per request. A 10-minute YouTube narration is about 9,850 characters, so 591 credits. It runs only on purchased credits, which never expire and come with no subscription; the free monthly credits you get with Google are for the web generator only.
Below, the cost per minute of audio on each service's cheapest plan, as published in September 2026. Voizum is shown with the Máxi pack, the one with the best price per credit. Figures are in euros for every brand so they can be compared.
| Service | Per minute of audio (entry plan) | API included in that plan | Unused credits |
|---|---|---|---|
| Voizum | €0.013 | Yes | Never expire |
| MiniMax | €0.043 | No, billed separately | Expire after 2 months |
| Fish Audio | €0.067 | No, billed separately | Reset every month |
| ElevenLabs | €0.171 | Yes | Lost when you cancel |
| Narakeet | €0.173 | Yes | Never expire |
| TTSMaker | €0.040 | No, billed separately | Reset every month |
Don't want to write code?
If what you want is to generate audio while you chat with Claude or ChatGPT, you do not need the API: the Voizum connector adds generate and list-voices tools to the assistant, signs in with your account, and uses the same rules as the web, free monthly credits included. Setup takes a couple of minutes from the Connect page.
And if you only need a web page to read text aloud in the visitor's own browser, the Web Speech API built into browsers is enough. It's free, but it sounds robotic, changes from one device to another and gives you no file. A TTS API is worth paying for once you need a natural voice in an MP3 you can publish.