Voizum

Guides

Text to speech API: guide with Python and JS examples

The Voizum team · Updated · 10 min read

A text to speech API turns text into an audio file from your own code or automation, with no one clicking buttons. This guide covers how to pick one for your use case, then walks through a complete, working integration: create the job, check its status, download the MP3. It also covers long scripts, batches, using your own cloned voice and plugging it into n8n, Make or Zapier.

What a text to speech API does and when you need one

A TTS API is an HTTP endpoint: you send text and a voice, and you get audio back. It makes sense as soon as audio becomes part of a process rather than a one-off task. If you narrate one video a month, a web editor is faster. If you narrate one a day, or your product needs to speak, an API saves the copy-paste-download loop every time.

Where we see it used most:

  • YouTube and faceless channels: the script comes out of a writing step (often an AI model) and goes straight to narration, then to the video editor.
  • E-learning and courses: hundreds of short lessons regenerated whenever the text changes.
  • Apps and products: notifications, onboarding, accessibility, read-aloud buttons for articles.
  • Batches: product descriptions, audio versions of a blog, phone menus, multilingual versions of the same script.
  • No-code automations: n8n, Make or Zapier flows that turn a new row in a spreadsheet into an MP3.

How to choose a text to speech API

Start from what you are building, not from the voice demo. The question that splits the market is latency versus length: a voice assistant needs the first syllable in a few hundred milliseconds and streams audio as it is generated, while a narrator needs a 40-minute script to come out as one clean file. Few APIs are good at both, and the cheapest per character are often the least natural.

CriterionWhy it mattersWhat to check
Latency vs. long audioReal-time agents need streaming; narration needs long files without seamsIs there streaming? What is the max text per request?
LanguagesEach language has its own set of voicesDoes the same voice speak every language you need?
Custom and cloned voicesA recognizable voice is a brand assetCan you use your own cloned voice through the API? Is there a cap on how many?
Price per characterIt is how almost every provider ends up billingPrice per 1,000 characters, minimum per request, whether the API costs extra
Unused creditsMonthly plans reset; you pay for what you did not useDo credits roll over, expire or vanish when you cancel?
LimitsRequests per minute and jobs in progress shape your architectureRate-limit headers, Retry-After, max concurrent jobs
Output and storageYou will store, edit or stream the fileFormat (MP3, WAV, PCM), how long the provider keeps it

Get your key and your voice_id

Everything is plain REST with JSON, base URL https://voizum.com/api/v1, and authentication with a Bearer key that starts with sk_voizum_. The flow is always the same: create a job, check its status, download the MP3.

  1. 1Sign in to Voizum with Google and buy any credit pack: the API key unlocks with both. The API spends purchased credits, never the free monthly ones.
  2. 2Open the API page and generate your key. It is shown once, so store it right away in an environment variable (VOIZUM_API_KEY) or your secret manager.
  3. 3Add a voice to your library: clone your own or save one from the public library with Add. The API only works with voices in your library.
  4. 4List them with GET /voices. Each item has id, name and language; the id is the voice_id you send when generating. You can also copy it in the web: Voices, the ⋯ menu, API ID.
  5. 5Optional: before a big run, GET /status (public, no key) tells you the service is up, and GET /account shows your credit balance and limits.

Code examples: curl, Python and JavaScript

The examples below work as is once you replace YOUR_VOICE_ID and JOB_ID. The audio URL returned by the status call needs the same Authorization header and redirects to the file, so let your HTTP client follow redirects.

curl

Create the job (202 with job_id and eta_seconds), poll it, then download:

# 1. Create the job
curl -X POST https://voizum.com/api/v1/tts \
  -H "Authorization: Bearer $VOIZUM_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: video-042-intro" \
  -d '{"text": "Welcome back to the channel.", "voice_id": "YOUR_VOICE_ID"}'

# 2. Check its status (repeat until "done")
curl https://voizum.com/api/v1/tts/JOB_ID -H "Authorization: Bearer $VOIZUM_API_KEY"

# 3. Download the MP3
curl -L -o narration.mp3 https://voizum.com/api/v1/tts/JOB_ID/audio -H "Authorization: Bearer $VOIZUM_API_KEY"

Python (requests)

The same flow, waiting while the job is queued or processing:

import os, time, requests

API = "https://voizum.com/api/v1"
H = {"Authorization": "Bearer " + os.environ["VOIZUM_API_KEY"]}

# 1. Create the job
with open("script.txt", encoding="utf-8") as f:
    texto = f.read()
r = requests.post(API + "/tts", headers=H, json={"text": texto, "voice_id": "YOUR_VOICE_ID"})
r.raise_for_status()
job = r.json()["job_id"]

# 2. Wait while it is queued or processing
while True:
    s = requests.get(API + "/tts/" + job, headers=H).json()
    if s["status"] not in ("queued", "processing"):
        break
    time.sleep(5)
if s["status"] != "done":
    raise RuntimeError(s.get("error", s["status"]))

# 3. Download the MP3
audio = requests.get(s["audio_url"], headers=H)
audio.raise_for_status()
with open("narration.mp3", "wb") as f:
    f.write(audio.content)

JavaScript (Node 18+, ES modules)

Native fetch, no SDK needed:

import fs from "node:fs";

const API = "https://voizum.com/api/v1";
const H = { Authorization: "Bearer " + process.env.VOIZUM_API_KEY, "Content-Type": "application/json" };

// 1. Create the job
const crear = await fetch(API + "/tts", {
  method: "POST",
  headers: H,
  body: JSON.stringify({ text: "Welcome back to the channel.", voice_id: "YOUR_VOICE_ID" }),
});
if (!crear.ok) throw new Error(await crear.text());
const { job_id } = await crear.json();

// 2. Wait while it is queued or processing
let s;
do {
  await new Promise((r) => setTimeout(r, 5000));
  s = await (await fetch(API + "/tts/" + job_id, { headers: H })).json();
} while (s.status === "queued" || s.status === "processing");
if (s.status !== "done") throw new Error(s.error ?? s.status);

// 3. Download the MP3
const audio = await fetch(s.audio_url, { headers: H });
fs.writeFileSync("narration.mp3", Buffer.from(await audio.arrayBuffer()));

Optional parameters

speed (0.5 to 2, default 1), pause_ms (pause between sentences, 0 to 800 ms), language ("auto" or es, en, de, fr, pt, it, ru; by default it is detected from the text) and loudness_normalization (true to make every voice come out at the same level). Status values are queued, processing, done, error and canceled. The full reference is at https://voizum.com/docs, with an OpenAPI spec at https://voizum.com/openapi.json and a plain-text version for AI coding assistants at https://voizum.com/llms.txt.

Long scripts, batches and dialogue

One request takes up to 600,000 characters, roughly 10 hours of speech, and comes back as a single MP3. You do not need to split a long script into chunks and stitch them yourself: the job does that internally and the status call shows progress (parts_done, parts_total, percent).

For many short pieces, use POST /tts/batch: 100 items maximum, 1,200,000 characters in total (about 20 hours). Each item gets its own job_id and its own file, but the 100-credit minimum is charged once for the whole batch instead of once per request. Eight short app prompts cost 800 credits sent one by one and 100 in a batch.

For a conversation in one file (a podcast intro, an ad with two voices) use POST /tts/dialogue with an ordered turns array, up to 5 different voices. It is billed like a normal audio, by the characters of the whole script.

Using your cloned voice through the API

Cloning happens once in the web and costs 199 credits; after that the voice is just another voice_id and every generation is billed like any other voice. There is no clone endpoint in the API on purpose: the recording is the step where you want to listen before committing. Record 15 to 60 seconds of only your voice, cut on a pause rather than mid-word, and talk the way you want the narration to sound, because the intonation is copied too.

Your cloned voices can speak all seven languages, so one English recording also narrates the Spanish, German or French version of your channel. Clone only your own voice or one you have permission to use.

Connecting it to n8n, Make or Zapier

No automation platform needs a special Voizum module: all three have a generic HTTP step, and the API is three plain calls. The only design choice is how you wait for the audio: poll in a loop, or let the API call you back with webhook_url.

The webhook option (POST /tts only) sends a POST with event, id, status (completed or error), voice_id, duration_seconds and credits when the job ends, retried up to three times. It must be a public https URL, so a self-hosted n8n on localhost cannot receive it: poll instead.

n8n

HTTP Request node with Header Auth credentials (Authorization, value Bearer sk_voizum_…), then a loop for the wait:

  • HTTP Request: POST https://voizum.com/api/v1/tts with a JSON body containing text and voice_id.
  • Wait node for a few seconds, then HTTP Request: GET /tts/{{ $json.job_id }}.
  • If node: status is done, continue; queued or processing, back to Wait; error, stop the run.
  • HTTP Request on audio_url with the same credentials and the response format set to File; the binary goes on to Drive, S3 or your video step.

Make and Zapier

In Make, the HTTP app's Make a request module sends the POST with the Authorization header. In Zapier, Webhooks by Zapier does the same with a POST or Custom Request action. In both, the cleanest wait is a second scenario or Zap that starts from a custom webhook (Make) or a Catch Hook trigger (Zapier): pass that URL as webhook_url, and when it fires, fetch GET /tts/{id} and download audio_url with your key.

Errors and good practices in production

Every error has the same shape, { error: { code, message } }. Program against code, which is stable, never against message, which is human text and can change.

  • Send an Idempotency-Key on every POST /tts (your video or row id works). If a network error makes you retry, you get the same job back with a 200 and are not charged twice.
  • Save the job_id as soon as you get it. If your process crashes, GET /tts lists your recent jobs with a cursor, and you can filter with ?status=done.
  • Poll every few seconds, or sleep for eta_seconds, rather than in a tight loop.
  • Download and keep your own copy: the MP3 is kept for 4 days, then the file is deleted.
  • Made a mistake? DELETE /tts/{id} cancels a job that is still queued and refunds it in full. Once it is processing it cannot be cancelled. If a job ends in error, its credits are returned automatically.
  • Never call the API from the browser or a mobile app: the key would be visible to anyone. Call it from your server, and have your frontend talk to your server.
  • Rotate the key if it may have leaked. Generating a new one revokes the old one immediately, so update the environment variable straight after.
codeHTTPWhat to do
no_autorizado401Missing, wrong or revoked key
saldo_insuficiente402Not enough purchased credits; the response says how many you need
voz_no_encontrada404That voice_id is not in your library: add it first
texto_demasiado_largo400Over 600,000 characters: split it
limite_peticiones429Wait the seconds in Retry-After, then retry
demasiados_en_cola429Too many jobs in progress: let some finish
servicio_no_disponible503Retry later with backoff
Limits: 20 new audios per minute and 300 requests per minute in total, per key. Every response carries X-RateLimit-Limit and X-RateLimit-Remaining.

How much does a text to speech API cost?

The API spends the same credits as the web: 60 credits per 1,000 characters (about 60 per minute of audio), with a minimum of 100 per request. A 10-minute YouTube narration is about 9,850 characters, so 591 credits. It runs only on purchased credits, which never expire and come with no subscription; the free monthly credits you get with Google are for the web generator only.

Below, the cost per minute of audio on each service's cheapest plan, as published in September 2026. Voizum is shown with the Máxi pack, the one with the best price per credit. Figures are in euros for every brand so they can be compared.

ServicePer minute of audio (entry plan)API included in that planUnused credits
Voizum€0.013YesNever expire
MiniMax€0.043No, billed separatelyExpire after 2 months
Fish Audio€0.067No, billed separatelyReset every month
ElevenLabs€0.171YesLost when you cancel
Narakeet€0.173YesNever expire
TTSMaker€0.040No, billed separatelyReset every month
Where the API is billed separately, its per-minute price can differ from the plan's. Cheaper standard voices from the big cloud providers exist; they sound less natural and do not clone voices.

Don't want to write code?

If what you want is to generate audio while you chat with Claude or ChatGPT, you do not need the API: the Voizum connector adds generate and list-voices tools to the assistant, signs in with your account, and uses the same rules as the web, free monthly credits included. Setup takes a couple of minutes from the Connect page.

And if you only need a web page to read text aloud in the visitor's own browser, the Web Speech API built into browsers is enough. It's free, but it sounds robotic, changes from one device to another and gives you no file. A TTS API is worth paying for once you need a natural voice in an MP3 you can publish.

Sources

  1. n8n docs: HTTP Request node
  2. n8n docs: HTTP Request credentials (Header Auth)
  3. Make developer docs: making requests
  4. Zapier help: send webhooks in Zaps
  5. IETF draft: the Idempotency-Key HTTP header field
  6. MDN: Retry-After header
  7. MDN: Web Speech API
  8. OWASP: Secrets Management Cheat Sheet

You might also find these useful

Frequently asked questions

Is there a free text to speech API?
Yes, with limits. Browsers include the free Web Speech API, which reads text aloud on the user's device but returns no file and sounds synthetic. Open-source models are free if you host them on your own GPU. Voizum's API runs on purchased credits; the free monthly credits only work in the web generator.
How do I use a text to speech API in Python?
Send a POST with your text and voice_id using requests, save the job_id it returns, poll the status every few seconds until it is done, and download the audio_url with the same Authorization header. The Python example in this guide does exactly that.
Can I use my own cloned voice through the API?
Yes. Clone it once in the web (199 credits), and its id appears in GET /voices like any other voice in your library. From then on you use it as voice_id and it can speak all seven supported languages.
How long can the text be in one request?
Up to 600,000 characters per request, about 10 hours of audio, delivered as one MP3. For many separate pieces, a batch request takes up to 100 items with the minimum charged only once.
Is the Voizum API good for real-time voice agents?
No. It is asynchronous and has no streaming: you create a job and collect the file when it is ready. It is built for narration, courses, batches and automations; for live conversation, pick a streaming API.
Can I use the generated audio commercially?
Yes, on YouTube, ads, courses or inside your product. What is not allowed is impersonating someone or using a voice you have no permission to clone, and illegal content.
How do I connect a TTS API to n8n?
Use the HTTP Request node with Header Auth for the Bearer key: one node creates the job, a Wait and an If node loop on the status, and a last HTTP Request downloads audio_url as a file. If your n8n has a public https address, you can pass a webhook URL instead of looping.

More guides

Type your text and hear it in a real voice

The API spends credits from a pack (60 per 1,000 characters) and needs a Google sign-in. Packs from $6.99, no subscription.