Guides
How to convert text to speech with AI, step by step
The Voizum team · Updated · 11 min read
Text to speech turns written text into spoken audio. A few years ago it sounded like an automated phone menu. Today's neural voices take a breath between sentences and lift at the end of a question. This guide covers how to choose a tool, with per-minute prices side by side, how to write numbers and acronyms so they don't trip the voice up, how long your audio will run, what bitrate each platform wants and what you can legally do with the result.
What text to speech is, and why it stopped sounding robotic
You already use it when your maps app tells you to take the next exit. Text to speech (TTS) is any system that reads written text aloud and hands you the sound.
Older voices were built by stitching together snippets of recordings or computing the sound from fixed rules. Hence the flat tone and choppy syllables: the system didn't understand the sentence, it just sounded it out.
Neural voices learn to speak from thousands of hours of real people and generate the waveform from scratch. The jump was measured with WaveNet. In Google's tests, listeners rated its US English voices 4.1 out of 5, more than 20% better than standard voices, closing the gap with human speech by over 70%. Since then, the model decides where to breathe and which word to stress.
So whether a tool sounds human is no longer the question. The real differences are more mundane: how much text it takes, whether you can download the MP3, whether you can use it commercially and what each minute costs.
What should you check before picking a tool?
Before you paste your script into the first free site that shows up, check these six things. You find out about the last four too late: the video's already edited, and it turns out you can't download the audio or can't monetize it.
| What to check | How to check it | Why it matters |
|---|---|---|
| Human-sounding voice | Listen to the voice reading YOUR text, not just the demo | A 10-second demo is picked to impress; a paragraph of yours with numbers and names tells the truth |
| Languages and accents | Native voices for each language, not an American voice reading Spanish | A foreign accent gives it away in the first sentence |
| Text limit | Characters per file and per month | If the free plan stops at a few thousand characters, you end up splitting the script and stitching the files back together |
| MP3 download | You get a file, not just playback in the browser | Without a file you can't drop it into a video editor or upload it anywhere |
| Commercial license | Whether the free tier allows monetized use and whether you must credit the tool | Some tools only allow commercial use on paid plans |
| Price per minute | Plan price divided by the minutes of audio it gives you | Every service defines a "credit" differently; minutes are the only fair comparison |
What does a minute of AI voice cost?
We compare per minute of audio, using the cheapest plan you can actually buy from each service. Where a price is published in characters, we convert it at our measured average of 1,000 characters per minute.
Watch the last column. On a subscription, what you don't use in a month is usually gone. With a one-off purchase, it stays.
| Service | Entry plan | Minutes | Price per minute | Unused balance |
|---|---|---|---|---|
| Voizum | Packs from $6.99 to $54.99, one-off | ≈ 365 – 3,723 | $0.015 – $0.019 (in Europe, €0.013 – €0.016) | Never expire |
| MiniMax | $5.00 one-off | ≈ 102 | $0.049 | Expire after 2 months |
| Fish Audio | €13.49 /month | ≈ 200 | €0.067 | Lost every month |
| ElevenLabs | $6.00 /month | ≈ 30 | $0.197 | Lost when you cancel |
| Narakeet | $6.00 one-off | ≈ 30 | $0.200 | Never expire |
| TTSMaker | $13.99 /month | ≈ 305 | $0.046 | Lost every month |
How to convert text to speech on Voizum, step by step
Here's how to turn text into a downloadable MP3 with nothing to install:
- 1Open the generator. Paste your text or upload a document (.txt, .md, .srt, .pdf, .docx). A single file can hold up to 600,000 characters, roughly 10 hours of speech.
- 2Pick one of 300+ voices, filtering by language, accent, gender or style. Play the sample first, and if you're torn between two, generate one paragraph with each.
- 3The language is detected from your text, in any of the 7 supported: English, Spanish, German, French, Portuguese, Italian and Russian.
- 4Adjust the speed (0.5× to 2×), the pause between sentences (up to 800 ms) and the volume. You can preview the settings before generating.
- 5Click Generate. You see the cost up front: 60 credits per 1,000 characters (about 60 per minute of audio), with a minimum of 100 per file.
- 6Listen back and download the MP3. Save it to your device: files stay in your history for 4 days.
Write it the way it sounds
The voice reads exactly what's on the page, punctuation included. Punctuation is how you direct it, and anything that's ambiguous on paper, like a date or an abbreviation, can come out in a way you didn't expect.
| If you write | Write this instead | Why |
|---|---|---|
| 10/2/2026 | October 2, 2026 | Half the world reads that as February 10; the voice can't know which you mean |
| St. Louis / Main St. | Saint Louis / Main Street | "St." can be Saint or Street, and gets guessed |
| $4.99/mo | four ninety-nine a month | Symbols and slashes can be read literally or skipped |
| 10-15 min | ten to fifteen minutes | The dash may be read as "minus" or dropped |
| voizum.com/guides | voizum dot com slash guides | URLs come out unpredictably, and nobody types one they only heard |
| NASA, SEO, CEO | NASA, S-E-O, C-E-O | Acronyms said as words are fine; spell out the ones you say letter by letter |
| (see footnote 3) | Cut it | Anything that only works on paper still gets read aloud |
Punctuation is your direction
A period makes a pause and closes the sentence; a comma, a short pause; a question mark lifts the ending. An ellipsis stretches the pause. If a line sounds rushed, it's almost always missing a comma.
For a longer silence, such as between sections, put a pause tag between two sentences: <#1.5#> leaves a second and a half (up to 10 seconds). Pause tags aren't read out and aren't charged.
Sentences built for the ear
Listeners can't glance back. One idea per sentence, with the subject close to the verb. And an old voiceover trick: a short sentence after a long one grabs attention without raising the volume.
- Read your script out loud first: wherever you run out of breath, split the sentence.
- Turn bullet points into spoken sequence: "First…, then…, and finally…".
- Repeat the subject after a couple of sentences: in audio, "it" and "this" get lost.
Mixing languages
Each file is generated in one language: whichever dominates the text. A word or a short quote in another language is read with the main language's voice and accent, the way a native speaker would. If your script has whole paragraphs in another language, generate those separately with a voice for that language and join the files.
Which voice to choose for each use
The voice shapes the result more than any setting. Think about who's listening and for how long: a high-energy voice sells a 30-second ad and wears people out over a six-hour audiobook.
| Use | What to look for | Pace | Practical tip |
|---|---|---|---|
| YouTube narration and documentaries | Warm narrator voice with some range | Normal | Test it on your opening hook: that's where viewers decide to stay |
| Ads and promos | Energetic, friendly voice | A bit faster | Short sentences, and the call to action alone at the end |
| Audiobooks and stories | Calm voice that doesn't tire the ear | Normal or a bit slower | Listen for 5 minutes straight, not 10 seconds: fatigue shows up over time |
| Training and tutorials | Clear, neutral voice | A bit slower | A long pause after each step gives people time to do it |
| News and recaps | Newsreader-style voice | Normal | Write figures the way they're said |
| Listening to your own notes | Whatever you find pleasant | Faster | With text you know, you can speed it up without losing the thread |
How long your audio will be
On average, one minute of audio is about 1,000 characters including spaces. But the voice matters more than the text: the same script runs almost 50% longer with the slowest voice we've measured than with the fastest.
| Characters | Fast voice | Average | Slow voice |
|---|---|---|---|
| 1,000 (A short ad or an intro) | 47 s | 1 min 1 s | 1 min 9 s |
| 5,000 (A blog post) | 3 min 53 s | 5 min 5 s | 5 min 45 s |
| 15,000 (A long YouTube video script) | 11 min 38 s | 15 min 14 s | 17 min 14 s |
| 60,000 (A long chapter or a short course) | 46 min 31 s | 1 h 1 min | 1 h 9 min |
Audio format and quality: MP3 and bitrate
Voizum delivers mono MP3 at 96 kbps, which opens in any player or video editor. A voice only needs one channel, and at that bitrate there's no audible loss in speech. Each minute is about 0.72 MB; an hour, roughly 43 MB.
If a platform asks for more, exporting at 192 kbps meets the spec, but it still sounds like the MP3 you started with.
| Destination | What it recommends | Source |
|---|---|---|
| YouTube | AAC-LC or Opus audio at 48 kHz; 128 kbps for mono | YouTube Help, recommended upload encoding |
| Apple Podcasts | MP3 or AAC at 128–256 kbps, loudness around −16 LKFS | Apple Podcasts for Creators |
| Audible audiobook (ACX) | MP3 at 192 kbps or higher, 44.1 kHz, −23 to −18 dB RMS | ACX submission requirements |
| Listening on your phone or sharing | The MP3 as is | No requirement |
Copyright and commercial use
On Voizum, the audio you generate is yours to use commercially: monetized videos, ads, courses or podcasts, with no credit to the tool required. Two limits apply: you can't use it to impersonate someone or for illegal content, and you can't clone someone else's voice without their permission.
And here's what the platforms ask:
- YouTube doesn't require you to disclose an AI voiceover or dub. It does require labeling realistic content that makes a real person appear to say something they didn't.
- To earn from ads, YouTube excludes "inauthentic" content: mass-produced, repetitive videos with no original input. The AI voice is fine; what gets a channel turned down is a recycled script or a template repeated a thousand times.
- Audible, through ACX, requires audiobooks to be narrated by a human unless they're part of its own authorized synthetic voice programs. If Audible is your target, confirm this before producing the whole book.
Free text to speech: what each option gives you
To hear a text once, you don't need a website at all. Microsoft Edge has a free Read Aloud with natural voices, and your operating system has a built-in reader. The catch: they read live and don't give you a file.
If you need the MP3, sign in with Google and you get 600 free credits every month, about 10 minutes of audio, to use with any voice in the library. They renew automatically, don't roll over and need no card.
If you need more, credits are bought with no subscription and never expire. The free monthly credits are for generating audio on the website; voice cloning, the API and batch generation use purchased credits.
Common mistakes when converting text to audio
The ones you hear first:
- Pasting text without cleaning it up: stray line breaks, page numbers or leftover formatting create odd pauses or get read out.
- Picking a voice from its 10-second demo without testing a real paragraph of your script.
- Leaving acronyms, abbreviations and ambiguous dates (10/2/2026) the way a document shows them instead of the way people say them.
- Four-line sentences: fine on paper, but in audio the listener loses the subject.
- Generating a long script in one go without listening to a minute first: if the voice or the speed is off, you pay for all of it.
- Not downloading the MP3: files are removed from your history after 4 days.
Sources
- Google Cloud: Introducing Cloud Text-to-Speech powered by DeepMind WaveNet (MOS 4.1)
- YouTube Help: Recommended upload encoding settings
- Apple Podcasts for Creators: Audio requirements
- ACX: Audio submission requirements
- YouTube Help: Disclosing altered or synthetic content
- YouTube Help: Channel monetization policies
- Microsoft Support: Use reading mode and Read Aloud in Microsoft Edge
- ElevenLabs: Pricing
- Fish Audio: Plans
- MiniMax Audio
- Narakeet: Pricing
- TTSMaker: Pricing