Voizum

Guides

How to convert text to speech with AI, step by step

The Voizum team · Updated · 11 min read

Text to speech turns written text into spoken audio. A few years ago it sounded like an automated phone menu. Today's neural voices take a breath between sentences and lift at the end of a question. This guide covers how to choose a tool, with per-minute prices side by side, how to write numbers and acronyms so they don't trip the voice up, how long your audio will run, what bitrate each platform wants and what you can legally do with the result.

What text to speech is, and why it stopped sounding robotic

You already use it when your maps app tells you to take the next exit. Text to speech (TTS) is any system that reads written text aloud and hands you the sound.

Older voices were built by stitching together snippets of recordings or computing the sound from fixed rules. Hence the flat tone and choppy syllables: the system didn't understand the sentence, it just sounded it out.

Neural voices learn to speak from thousands of hours of real people and generate the waveform from scratch. The jump was measured with WaveNet. In Google's tests, listeners rated its US English voices 4.1 out of 5, more than 20% better than standard voices, closing the gap with human speech by over 70%. Since then, the model decides where to breathe and which word to stress.

So whether a tool sounds human is no longer the question. The real differences are more mundane: how much text it takes, whether you can download the MP3, whether you can use it commercially and what each minute costs.

What should you check before picking a tool?

Before you paste your script into the first free site that shows up, check these six things. You find out about the last four too late: the video's already edited, and it turns out you can't download the audio or can't monetize it.

What to checkHow to check itWhy it matters
Human-sounding voiceListen to the voice reading YOUR text, not just the demoA 10-second demo is picked to impress; a paragraph of yours with numbers and names tells the truth
Languages and accentsNative voices for each language, not an American voice reading SpanishA foreign accent gives it away in the first sentence
Text limitCharacters per file and per monthIf the free plan stops at a few thousand characters, you end up splitting the script and stitching the files back together
MP3 downloadYou get a file, not just playback in the browserWithout a file you can't drop it into a video editor or upload it anywhere
Commercial licenseWhether the free tier allows monetized use and whether you must credit the toolSome tools only allow commercial use on paid plans
Price per minutePlan price divided by the minutes of audio it gives youEvery service defines a "credit" differently; minutes are the only fair comparison

What does a minute of AI voice cost?

We compare per minute of audio, using the cheapest plan you can actually buy from each service. Where a price is published in characters, we convert it at our measured average of 1,000 characters per minute.

Watch the last column. On a subscription, what you don't use in a month is usually gone. With a one-off purchase, it stays.

ServiceEntry planMinutesPrice per minuteUnused balance
VoizumPacks from $6.99 to $54.99, one-off≈ 365 – 3,723$0.015 – $0.019 (in Europe, €0.013 – €0.016)Never expire
MiniMax$5.00 one-off≈ 102$0.049Expire after 2 months
Fish Audio€13.49 /month≈ 200€0.067Lost every month
ElevenLabs$6.00 /month≈ 30$0.197Lost when you cancel
Narakeet$6.00 one-off≈ 30$0.200Never expire
TTSMaker$13.99 /month≈ 305$0.046Lost every month
Prices from each brand's official site: entry plan on monthly billing (never the annual plan split into months), in the currency they publish, checked in September and October 2026. Prices change, so check before you buy.

How to convert text to speech on Voizum, step by step

Here's how to turn text into a downloadable MP3 with nothing to install:

  1. 1Open the generator. Paste your text or upload a document (.txt, .md, .srt, .pdf, .docx). A single file can hold up to 600,000 characters, roughly 10 hours of speech.
  2. 2Pick one of 300+ voices, filtering by language, accent, gender or style. Play the sample first, and if you're torn between two, generate one paragraph with each.
  3. 3The language is detected from your text, in any of the 7 supported: English, Spanish, German, French, Portuguese, Italian and Russian.
  4. 4Adjust the speed (0.5× to 2×), the pause between sentences (up to 800 ms) and the volume. You can preview the settings before generating.
  5. 5Click Generate. You see the cost up front: 60 credits per 1,000 characters (about 60 per minute of audio), with a minimum of 100 per file.
  6. 6Listen back and download the MP3. Save it to your device: files stay in your history for 4 days.

Write it the way it sounds

The voice reads exactly what's on the page, punctuation included. Punctuation is how you direct it, and anything that's ambiguous on paper, like a date or an abbreviation, can come out in a way you didn't expect.

If you writeWrite this insteadWhy
10/2/2026October 2, 2026Half the world reads that as February 10; the voice can't know which you mean
St. Louis / Main St.Saint Louis / Main Street"St." can be Saint or Street, and gets guessed
$4.99/mofour ninety-nine a monthSymbols and slashes can be read literally or skipped
10-15 minten to fifteen minutesThe dash may be read as "minus" or dropped
voizum.com/guidesvoizum dot com slash guidesURLs come out unpredictably, and nobody types one they only heard
NASA, SEO, CEONASA, S-E-O, C-E-OAcronyms said as words are fine; spell out the ones you say letter by letter
(see footnote 3)Cut itAnything that only works on paper still gets read aloud

Punctuation is your direction

A period makes a pause and closes the sentence; a comma, a short pause; a question mark lifts the ending. An ellipsis stretches the pause. If a line sounds rushed, it's almost always missing a comma.

For a longer silence, such as between sections, put a pause tag between two sentences: <#1.5#> leaves a second and a half (up to 10 seconds). Pause tags aren't read out and aren't charged.

Sentences built for the ear

Listeners can't glance back. One idea per sentence, with the subject close to the verb. And an old voiceover trick: a short sentence after a long one grabs attention without raising the volume.

  • Read your script out loud first: wherever you run out of breath, split the sentence.
  • Turn bullet points into spoken sequence: "First…, then…, and finally…".
  • Repeat the subject after a couple of sentences: in audio, "it" and "this" get lost.

Mixing languages

Each file is generated in one language: whichever dominates the text. A word or a short quote in another language is read with the main language's voice and accent, the way a native speaker would. If your script has whole paragraphs in another language, generate those separately with a voice for that language and join the files.

Which voice to choose for each use

The voice shapes the result more than any setting. Think about who's listening and for how long: a high-energy voice sells a 30-second ad and wears people out over a six-hour audiobook.

UseWhat to look forPacePractical tip
YouTube narration and documentariesWarm narrator voice with some rangeNormalTest it on your opening hook: that's where viewers decide to stay
Ads and promosEnergetic, friendly voiceA bit fasterShort sentences, and the call to action alone at the end
Audiobooks and storiesCalm voice that doesn't tire the earNormal or a bit slowerListen for 5 minutes straight, not 10 seconds: fatigue shows up over time
Training and tutorialsClear, neutral voiceA bit slowerA long pause after each step gives people time to do it
News and recapsNewsreader-style voiceNormalWrite figures the way they're said
Listening to your own notesWhatever you find pleasantFasterWith text you know, you can speed it up without losing the thread

How long your audio will be

On average, one minute of audio is about 1,000 characters including spaces. But the voice matters more than the text: the same script runs almost 50% longer with the slowest voice we've measured than with the fastest.

CharactersFast voiceAverageSlow voice
1,000 (A short ad or an intro)47 s1 min 1 s1 min 9 s
5,000 (A blog post)3 min 53 s5 min 5 s5 min 45 s
15,000 (A long YouTube video script)11 min 38 s15 min 14 s17 min 14 s
60,000 (A long chapter or a short course)46 min 31 s1 h 1 min1 h 9 min
Pace measured on real generated audio: 14.5 to 21.5 characters per second depending on the voice, at 1× speed.

Audio format and quality: MP3 and bitrate

Voizum delivers mono MP3 at 96 kbps, which opens in any player or video editor. A voice only needs one channel, and at that bitrate there's no audible loss in speech. Each minute is about 0.72 MB; an hour, roughly 43 MB.

If a platform asks for more, exporting at 192 kbps meets the spec, but it still sounds like the MP3 you started with.

DestinationWhat it recommendsSource
YouTubeAAC-LC or Opus audio at 48 kHz; 128 kbps for monoYouTube Help, recommended upload encoding
Apple PodcastsMP3 or AAC at 128–256 kbps, loudness around −16 LKFSApple Podcasts for Creators
Audible audiobook (ACX)MP3 at 192 kbps or higher, 44.1 kHz, −23 to −18 dB RMSACX submission requirements
Listening on your phone or sharingThe MP3 as isNo requirement

Free text to speech: what each option gives you

To hear a text once, you don't need a website at all. Microsoft Edge has a free Read Aloud with natural voices, and your operating system has a built-in reader. The catch: they read live and don't give you a file.

If you need the MP3, sign in with Google and you get 600 free credits every month, about 10 minutes of audio, to use with any voice in the library. They renew automatically, don't roll over and need no card.

If you need more, credits are bought with no subscription and never expire. The free monthly credits are for generating audio on the website; voice cloning, the API and batch generation use purchased credits.

Common mistakes when converting text to audio

The ones you hear first:

  • Pasting text without cleaning it up: stray line breaks, page numbers or leftover formatting create odd pauses or get read out.
  • Picking a voice from its 10-second demo without testing a real paragraph of your script.
  • Leaving acronyms, abbreviations and ambiguous dates (10/2/2026) the way a document shows them instead of the way people say them.
  • Four-line sentences: fine on paper, but in audio the listener loses the subject.
  • Generating a long script in one go without listening to a minute first: if the voice or the speed is off, you pay for all of it.
  • Not downloading the MP3: files are removed from your history after 4 days.

Sources

  1. Google Cloud: Introducing Cloud Text-to-Speech powered by DeepMind WaveNet (MOS 4.1)
  2. YouTube Help: Recommended upload encoding settings
  3. Apple Podcasts for Creators: Audio requirements
  4. ACX: Audio submission requirements
  5. YouTube Help: Disclosing altered or synthetic content
  6. YouTube Help: Channel monetization policies
  7. Microsoft Support: Use reading mode and Read Aloud in Microsoft Edge
  8. ElevenLabs: Pricing
  9. Fish Audio: Plans
  10. MiniMax Audio
  11. Narakeet: Pricing
  12. TTSMaker: Pricing

You might also find these useful

Frequently asked questions

Is there free, unlimited text to speech?
Only if you don't need a file: browser and operating system readers are free and unlimited, but they don't save the audio. Sites that generate free MP3s usually cap characters per file or per month. On Voizum, signing in with Google gets you 600 free credits a month, about 10 minutes of audio.
How do I download text to speech as an MP3?
Generate the audio, listen back and click download: you get an MP3 that opens in any player or video editor. Save it to your device, since files stay in your history for 4 days.
Can I get a male or female voice?
Yes. The library has 300+ voices, and you can filter by gender, language, accent and style (narrator, ads, news). Play each sample before you generate.
Which text to speech sounds most like a human voice?
Any modern neural voice beats the old robotic ones; the difference between tools now shows up in long scripts, numbers and names. Test each option with a real paragraph of your own text rather than its demo.
How much text can I convert at once?
Up to 600,000 characters per file, roughly 10 hours of speech. For reference, a 15-minute video script is around 15,000 characters.
Can I use AI text to speech on YouTube and get monetized?
Yes, commercial use is allowed. YouTube doesn't ask you to disclose an AI voiceover, but it won't monetize mass-produced content with no original input, so what counts is that your video adds something.
Can I use text to speech with my own voice?
Yes. You clone your voice from a 15 to 60 second recording and use it like any other voice. Cloning costs 199 credits and requires having bought credits at least once.
Which languages does it support?
English, Spanish, German, French, Portuguese, Italian and Russian, with native voices and several accents (American and British English, Latin American and European Spanish). The language is detected from your text.

More guides

Type your text and hear it in a real voice

Sign in with Google and get 600 free credits every month, about 10 minutes of audio.