Voizum

Guides

AI voiceover generator for videos: script to publish

The Voizum team · Updated · 10 min read

An AI voiceover takes four steps: script, voice, edit, publish. This guide puts numbers on each one: how many words fill a minute, which voice suits which video, how to sync the MP3 in CapCut, Premiere, DaVinci or PowerPoint, how loud the music should sit, and what YouTube and TikTok expect when the narration is synthetic.

How to make an AI voiceover video, start to finish

Script, then voice, then picture. Cut the footage first and you'll spend the afternoon stretching shots or chopping sentences to squeeze the narration in. Lock the voice and the picture lines up on the first pass.

Same workflow for a 40-second Short or an hour-long documentary:

  1. 1Write the whole script in a doc, for the ear rather than the page. Tips below.
  2. 2Estimate the runtime: about 150 words, or roughly 1,000 characters, per minute. An 8-minute video needs around 1,200 words.
  3. 3Paste the text into the generator or upload the file (txt, md, srt, pdf, docx) and pick a voice. Preview it with a line from your own script, not the sample sentence.
  4. 4Generate and listen to the result. If you want to change something, change it in the text and generate just that part again.
  5. 5Download the MP3 and import it into your editor on its own audio track, starting at zero.
  6. 6Cut the footage to the voice: a new idea, a new shot.
  7. 7Drop music on a separate track, pull it under the voice and turn on auto-ducking if your editor has it.
  8. 8Add captions. A large share of social video is watched with the sound off.
  9. 9Export, check the final loudness, and publish. Use the platform's AI disclosure setting when it applies.

Voiceover words per minute: how long should your script be?

The usual reference is about 150 words per minute: the National Center for Voice and Speech puts average American English speech there. Audiobook narration tends to sit near that pace too. Use it to plan, then check it against the actual voice you pick, because voices differ more than scripts do.

Character count is the more precise unit, since that's what any voice generator measures. One thousand characters, spaces included, is a good round number for a minute of audio.

Video lengthWords (at 150 wpm)Characters (round count)Real range by voiceCredits
1 min1501,000870 – 1,290100
3 min4503,0002,610 – 3,870180
8 min1,2008,0006,960 – 10,320480
15 min2,25015,00013,050 – 19,350900
30 min4,50030,00026,100 – 38,7001,800
1 hour9,00060,00052,200 – 77,4003,600
Credits at 60 per 1,000 characters, with a 100-credit minimum per audio. Sign in with Google and you get 600 free credits a month, about 10 minutes of voiceover.

How to write a voiceover script that sounds natural

Write for the ear. Viewers can't scroll back and reread, so every sentence has to land the first time.

  • Keep sentences under 20-25 words. If one needs two commas and a parenthesis, split it.
  • One idea per sentence, with the key point at the end, where the stress naturally falls.
  • Ordinary amounts, dates, times and prices read fine on their own. Acronyms and odd symbols are another story: spell them the way you want them said, like "N-A-S-A" versus "NASA", or "three times" instead of "3x".
  • Punctuation sets the pace. A period is a pause, a comma is a breath, and a script with no periods gets read in one go.
  • Don't read long lists aloud. Past four items, people forget; put them on screen and have the voice name the two that matter.
  • Read the script out loud yourself before generating. Wherever you stumble, the AI voice will sound strained too.

The hook for Shorts, Reels and TikTok

In short-form video, the first line decides whether people stay. Open with the payoff or the most surprising fact, never with "Hey guys, today we're going to". If you need an intro, it goes after the hook.

Custom pauses

Need an exact beat while a chart comes up? Type <#1.5#> between the two sentences, with however many seconds you want. The tag isn't read aloud or charged. Overall spacing between sentences is a separate setting, from tight to relaxed.

Choosing an AI narrator voice for each type of video

The voice sets the tone before the picture does. Choose with your script in front of you: a voice that's perfect for a tutorial can sound flat on an ad.

Video typeVoice that worksPaceTip
Documentary, history, true crimeDeep, measured narratorSlowWell-punctuated long sentences and a pause before each reveal
Tutorial, course, explainerClear, friendly voice, no theatricsMediumOne step per sentence; name on-screen elements exactly as they appear
Ad or promoUpbeat commercial readMedium-fastProduct and benefit in the first 5 seconds; call to action at the end, on its own
Kids' story, fictionWarm, expressive voiceSlowRaise expressiveness in the settings and pause between scenes
Shorts, Reels, TikTokConversational, like a friend telling you somethingFastHook in line one, zero intro
News recap, faceless explainerNeutral announcer voiceMediumDates and figures written as spoken

Library voice or your own cloned voice?

If your audience knows your voice, clone it: your videos keep sounding like you on days you don't record. For a faceless channel, a library voice skips that step and you can switch voices per series or per language.

Library voices

There are more than 300 voices in 7 languages: English, Spanish, German, French, Portuguese, Italian and Russian. Each has a sample to preview, and you can tune speed, volume and expressiveness, then save those settings on the voice so every video sounds the same.

Your cloned voice

You need 15 to 60 seconds of clean recording of one person. Cloning costs 199 credits, once, and is available after your first credit purchase. Delivery carries over: record the sample in a monotone and the narration comes out monotone. You can only clone your own voice or someone who has given you permission.

  • Record in a furnished room with no music or background noise.
  • End the clip on a pause, never mid-word, with a bit under a second of silence at the end.
  • Don't loop a short clip to make it longer: it makes the clone worse, not better.

How to add a voiceover in CapCut, Premiere Pro and DaVinci Resolve

The menus differ, the idea doesn't: the MP3 gets its own audio track, starts at zero, and the picture follows it.

CapCut (desktop and mobile)

The fastest route for TikTok, Reels and Shorts.

  • Click Import and add your footage and the voiceover MP3.
  • Drag the MP3 onto the timeline under the video track, aligned to the start.
  • Zoom into the timeline and Split the clips wherever the narration moves to a new idea.
  • Select the music and lower its volume in the audio panel; volume keyframes let you dip it only while the voice is talking.
  • Go to Captions > Auto captions, pick the voice's language and proofread the result.

Premiere Pro

  • File > Import (Ctrl/Cmd+I), then drag the MP3 to track A1 and the music to A2.
  • In the Essential Sound panel, tag the voice as Dialogue and the music as Music.
  • With the music selected, enable Ducking, set Duck Against to Dialogue, adjust how far it drops, and click Generate Keyframes.
  • Cut your shots on the video track to the sentences of the voice.

DaVinci Resolve

  • Import the MP3 into the Media Pool and drag it to an audio track on the Edit page.
  • Put the music on a different track.
  • On the Fairlight page, open Dynamics on the music track, enable the compressor and set its sidechain source to the voice track.
  • Lower the threshold until the music backs off when the voice comes in; a 3:1 to 5:1 ratio is a sensible starting point.

How to do a voiceover on PowerPoint with an AI voice

PowerPoint can record your own narration from a mic, and for an internal meeting that's often enough. For a deck that needs to sound polished, or one you'll update often, an AI voice saves re-recording: change the text, regenerate that slide's clip, done.

The cleanest setup is one audio clip per slide. "Chain audios" generates up to 100 clips in one go, each named, which suits a deck well.

  1. 1On each slide: Insert > Audio > Audio on My PC, and choose that slide's MP3.
  2. 2With the audio icon selected, open the Playback tab and set Start to Automatically. Check Hide During Show.
  3. 3On the Transitions tab, uncheck On Mouse Click and set After to the length of that clip.
  4. 4To turn it into a video: File > Export > Create a Video, choose Use Recorded Timings and Narrations, pick a resolution (up to 4K) and save as MP4.

How to make a voiceover sound better: music and loudness

A common editing rule of thumb: voice peaks between −12 and −6 dB, music around 20 dB lower while someone is talking. Music with vocals or heavy drums needs to sit lower still.

Auto-ducking pulls the music down under the voice and back up in the gaps. Premiere has it in Essential Sound, DaVinci does it with a sidechained compressor, and in CapCut you do it with volume keyframes.

Before you export, measure the integrated loudness of the full mix in LUFS. These are the targets the platforms publish themselves:

DestinationIntegrated loudnessMax true peakSource
Spotify (streaming audio)−14 LUFS−1 dBTPSpotify for Artists
Apple Podcasts−16 LKFS (±1 dB)−1 dBFSApple Podcasts for Creators
LUFS and LKFS measure the same thing. YouTube doesn't publish an official target; for video, a mix between −14 and −16 LUFS with peaks below −1 dB plays well almost everywhere.

Hours of AI narration in a single file

An hour-long documentary, a course or an audiobook can be generated in one pass: each audio takes up to 600,000 characters, about 10 hours of speech (7 to 12 depending on how fast the voice is). No splitting the script, and no risk of two halves sounding slightly different.

Chaining clips and the API use purchased credits; the monthly free credits are for generating on the website.

Commercial use, YouTube monetization and AI disclosure

Yes to monetized videos, ads and client work. No to impersonating someone or illegal content.

YouTube asks you to flag realistic content that's generated or meaningfully altered with AI, such as making a real person appear to say something they didn't. Its own list of exceptions includes cloning your own voice for voiceovers or dubs, and using AI for scripts or captions.

Monetization is a separate question. Since July 2025 YouTube calls mass-produced or repetitive uploads "inauthentic content", and that explicitly covers AI videos built from generic templates without the creator's own input. A channel of interchangeable videos is at risk with or without a synthetic voice. What keeps you safe is an original script with your own research and point of view.

TikTok asks you to turn on its AI-generated content label for realistic AI content, and bans misleading synthetic media about real people or events regardless of labeling.

Common AI voiceover mistakes

Most of these get fixed by rewriting the script before you generate.

  • Pasting a blog post or a PDF as is. It sounds like someone reading a brochure out loud.
  • Leaving "1/2" or "Inc." as symbols. They can come out in unexpected ways.
  • Choosing a voice from its demo line instead of your own script.
  • Music too loud, or music with lyrics under the narration.
  • Cutting the picture first and forcing the voice to fit.
  • Generating 20 minutes without testing one minute first.
  • Cranking expressiveness to the max to sound "more alive": overdoing it makes it sound less natural.
  • Skipping captions on social video.

Sources

  1. YouTube Help: Disclosing altered or synthetic content
  2. YouTube Help: YouTube channel monetization policies
  3. TikTok Support: AI-generated content
  4. National Center for Voice and Speech: Voice qualities and speaking rate
  5. Spotify for Artists: Loudness normalization
  6. Apple Podcasts for Creators: Audio requirements
  7. Microsoft Support: Add or delete audio in your PowerPoint presentation
  8. Microsoft Support: Turn your presentation into a video
  9. Adobe Help: Automatically duck audio in Premiere
  10. Larry Jordan: Automatically duck background music under dialog in DaVinci Resolve
  11. CapCut Help: How to recognise subtitles
  12. DIY Video Studio: Ideal audio levels for video
  13. Merriam-Webster: voice-over

You might also find these useful

Frequently asked questions

Is it voiceover or voice over?
All three spellings show up. Merriam-Webster lists it as "voice-over" with a hyphen, while "voiceover" as one word is common in the industry and in search. "Voice over" as two words is usually the verb: you voice over a video.
Can you monetize YouTube videos with an AI voiceover?
Yes. YouTube doesn't ban synthetic voices; what it won't monetize is mass-produced or repetitive content without original input. If the video shows something realistic that didn't happen, flag it as altered or synthetic when you upload.
Is there a free AI voiceover generator?
Yes. Sign in with Google and you get 600 free credits every month, about 10 minutes of voice, to generate on the website and download as MP3. For very short clips, CapCut's built-in voices also work.
How do I do a voiceover on PowerPoint?
Insert > Audio > Audio on My PC on each slide, set Start to Automatically on the Playback tab, and set each slide's timing on the Transitions tab. Then File > Export > Create a Video to get an MP4.
How many words is a 1-minute voiceover?
About 150 words, or roughly 1,000 characters. Depending on the voice, the real range runs from about 870 to 1,290 characters per minute, so test a minute before writing to a tight runtime.
How do I make my voiceover sound better?
Write short sentences and punctuate them, spell out acronyms the way they're said, keep the music about 20 dB under the voice with ducking, and aim for a final mix around −14 to −16 LUFS.
Can an AI narrator use my own voice?
Yes. Clone it once from 15 to 60 seconds of clean recording, then type any script and it's read in your voice, in any of the 7 languages.

More guides

Type your text and hear it in a real voice

Sign in with Google and get 600 free credits every month, about 10 minutes of audio.