Guides
AI voiceover generator for videos: script to publish
The Voizum team · Updated · 10 min read
An AI voiceover takes four steps: script, voice, edit, publish. This guide puts numbers on each one: how many words fill a minute, which voice suits which video, how to sync the MP3 in CapCut, Premiere, DaVinci or PowerPoint, how loud the music should sit, and what YouTube and TikTok expect when the narration is synthetic.
How to make an AI voiceover video, start to finish
Script, then voice, then picture. Cut the footage first and you'll spend the afternoon stretching shots or chopping sentences to squeeze the narration in. Lock the voice and the picture lines up on the first pass.
Same workflow for a 40-second Short or an hour-long documentary:
- 1Write the whole script in a doc, for the ear rather than the page. Tips below.
- 2Estimate the runtime: about 150 words, or roughly 1,000 characters, per minute. An 8-minute video needs around 1,200 words.
- 3Paste the text into the generator or upload the file (txt, md, srt, pdf, docx) and pick a voice. Preview it with a line from your own script, not the sample sentence.
- 4Generate and listen to the result. If you want to change something, change it in the text and generate just that part again.
- 5Download the MP3 and import it into your editor on its own audio track, starting at zero.
- 6Cut the footage to the voice: a new idea, a new shot.
- 7Drop music on a separate track, pull it under the voice and turn on auto-ducking if your editor has it.
- 8Add captions. A large share of social video is watched with the sound off.
- 9Export, check the final loudness, and publish. Use the platform's AI disclosure setting when it applies.
Voiceover words per minute: how long should your script be?
The usual reference is about 150 words per minute: the National Center for Voice and Speech puts average American English speech there. Audiobook narration tends to sit near that pace too. Use it to plan, then check it against the actual voice you pick, because voices differ more than scripts do.
Character count is the more precise unit, since that's what any voice generator measures. One thousand characters, spaces included, is a good round number for a minute of audio.
| Video length | Words (at 150 wpm) | Characters (round count) | Real range by voice | Credits |
|---|---|---|---|---|
| 1 min | 150 | 1,000 | 870 – 1,290 | 100 |
| 3 min | 450 | 3,000 | 2,610 – 3,870 | 180 |
| 8 min | 1,200 | 8,000 | 6,960 – 10,320 | 480 |
| 15 min | 2,250 | 15,000 | 13,050 – 19,350 | 900 |
| 30 min | 4,500 | 30,000 | 26,100 – 38,700 | 1,800 |
| 1 hour | 9,000 | 60,000 | 52,200 – 77,400 | 3,600 |
How to write a voiceover script that sounds natural
Write for the ear. Viewers can't scroll back and reread, so every sentence has to land the first time.
- Keep sentences under 20-25 words. If one needs two commas and a parenthesis, split it.
- One idea per sentence, with the key point at the end, where the stress naturally falls.
- Ordinary amounts, dates, times and prices read fine on their own. Acronyms and odd symbols are another story: spell them the way you want them said, like "N-A-S-A" versus "NASA", or "three times" instead of "3x".
- Punctuation sets the pace. A period is a pause, a comma is a breath, and a script with no periods gets read in one go.
- Don't read long lists aloud. Past four items, people forget; put them on screen and have the voice name the two that matter.
- Read the script out loud yourself before generating. Wherever you stumble, the AI voice will sound strained too.
The hook for Shorts, Reels and TikTok
In short-form video, the first line decides whether people stay. Open with the payoff or the most surprising fact, never with "Hey guys, today we're going to". If you need an intro, it goes after the hook.
Custom pauses
Need an exact beat while a chart comes up? Type <#1.5#> between the two sentences, with however many seconds you want. The tag isn't read aloud or charged. Overall spacing between sentences is a separate setting, from tight to relaxed.
Choosing an AI narrator voice for each type of video
The voice sets the tone before the picture does. Choose with your script in front of you: a voice that's perfect for a tutorial can sound flat on an ad.
| Video type | Voice that works | Pace | Tip |
|---|---|---|---|
| Documentary, history, true crime | Deep, measured narrator | Slow | Well-punctuated long sentences and a pause before each reveal |
| Tutorial, course, explainer | Clear, friendly voice, no theatrics | Medium | One step per sentence; name on-screen elements exactly as they appear |
| Ad or promo | Upbeat commercial read | Medium-fast | Product and benefit in the first 5 seconds; call to action at the end, on its own |
| Kids' story, fiction | Warm, expressive voice | Slow | Raise expressiveness in the settings and pause between scenes |
| Shorts, Reels, TikTok | Conversational, like a friend telling you something | Fast | Hook in line one, zero intro |
| News recap, faceless explainer | Neutral announcer voice | Medium | Dates and figures written as spoken |
Library voice or your own cloned voice?
If your audience knows your voice, clone it: your videos keep sounding like you on days you don't record. For a faceless channel, a library voice skips that step and you can switch voices per series or per language.
Library voices
There are more than 300 voices in 7 languages: English, Spanish, German, French, Portuguese, Italian and Russian. Each has a sample to preview, and you can tune speed, volume and expressiveness, then save those settings on the voice so every video sounds the same.
Your cloned voice
You need 15 to 60 seconds of clean recording of one person. Cloning costs 199 credits, once, and is available after your first credit purchase. Delivery carries over: record the sample in a monotone and the narration comes out monotone. You can only clone your own voice or someone who has given you permission.
- Record in a furnished room with no music or background noise.
- End the clip on a pause, never mid-word, with a bit under a second of silence at the end.
- Don't loop a short clip to make it longer: it makes the clone worse, not better.
How to add a voiceover in CapCut, Premiere Pro and DaVinci Resolve
The menus differ, the idea doesn't: the MP3 gets its own audio track, starts at zero, and the picture follows it.
CapCut (desktop and mobile)
The fastest route for TikTok, Reels and Shorts.
- Click Import and add your footage and the voiceover MP3.
- Drag the MP3 onto the timeline under the video track, aligned to the start.
- Zoom into the timeline and Split the clips wherever the narration moves to a new idea.
- Select the music and lower its volume in the audio panel; volume keyframes let you dip it only while the voice is talking.
- Go to Captions > Auto captions, pick the voice's language and proofread the result.
Premiere Pro
- File > Import (Ctrl/Cmd+I), then drag the MP3 to track A1 and the music to A2.
- In the Essential Sound panel, tag the voice as Dialogue and the music as Music.
- With the music selected, enable Ducking, set Duck Against to Dialogue, adjust how far it drops, and click Generate Keyframes.
- Cut your shots on the video track to the sentences of the voice.
DaVinci Resolve
- Import the MP3 into the Media Pool and drag it to an audio track on the Edit page.
- Put the music on a different track.
- On the Fairlight page, open Dynamics on the music track, enable the compressor and set its sidechain source to the voice track.
- Lower the threshold until the music backs off when the voice comes in; a 3:1 to 5:1 ratio is a sensible starting point.
How to do a voiceover on PowerPoint with an AI voice
PowerPoint can record your own narration from a mic, and for an internal meeting that's often enough. For a deck that needs to sound polished, or one you'll update often, an AI voice saves re-recording: change the text, regenerate that slide's clip, done.
The cleanest setup is one audio clip per slide. "Chain audios" generates up to 100 clips in one go, each named, which suits a deck well.
- 1On each slide: Insert > Audio > Audio on My PC, and choose that slide's MP3.
- 2With the audio icon selected, open the Playback tab and set Start to Automatically. Check Hide During Show.
- 3On the Transitions tab, uncheck On Mouse Click and set After to the length of that clip.
- 4To turn it into a video: File > Export > Create a Video, choose Use Recorded Timings and Narrations, pick a resolution (up to 4K) and save as MP4.
How to make a voiceover sound better: music and loudness
A common editing rule of thumb: voice peaks between −12 and −6 dB, music around 20 dB lower while someone is talking. Music with vocals or heavy drums needs to sit lower still.
Auto-ducking pulls the music down under the voice and back up in the gaps. Premiere has it in Essential Sound, DaVinci does it with a sidechained compressor, and in CapCut you do it with volume keyframes.
Before you export, measure the integrated loudness of the full mix in LUFS. These are the targets the platforms publish themselves:
| Destination | Integrated loudness | Max true peak | Source |
|---|---|---|---|
| Spotify (streaming audio) | −14 LUFS | −1 dBTP | Spotify for Artists |
| Apple Podcasts | −16 LKFS (±1 dB) | −1 dBFS | Apple Podcasts for Creators |
Hours of AI narration in a single file
An hour-long documentary, a course or an audiobook can be generated in one pass: each audio takes up to 600,000 characters, about 10 hours of speech (7 to 12 depending on how fast the voice is). No splitting the script, and no risk of two halves sounding slightly different.
Chaining clips and the API use purchased credits; the monthly free credits are for generating on the website.
Commercial use, YouTube monetization and AI disclosure
Yes to monetized videos, ads and client work. No to impersonating someone or illegal content.
YouTube asks you to flag realistic content that's generated or meaningfully altered with AI, such as making a real person appear to say something they didn't. Its own list of exceptions includes cloning your own voice for voiceovers or dubs, and using AI for scripts or captions.
Monetization is a separate question. Since July 2025 YouTube calls mass-produced or repetitive uploads "inauthentic content", and that explicitly covers AI videos built from generic templates without the creator's own input. A channel of interchangeable videos is at risk with or without a synthetic voice. What keeps you safe is an original script with your own research and point of view.
TikTok asks you to turn on its AI-generated content label for realistic AI content, and bans misleading synthetic media about real people or events regardless of labeling.
Common AI voiceover mistakes
Most of these get fixed by rewriting the script before you generate.
- Pasting a blog post or a PDF as is. It sounds like someone reading a brochure out loud.
- Leaving "1/2" or "Inc." as symbols. They can come out in unexpected ways.
- Choosing a voice from its demo line instead of your own script.
- Music too loud, or music with lyrics under the narration.
- Cutting the picture first and forcing the voice to fit.
- Generating 20 minutes without testing one minute first.
- Cranking expressiveness to the max to sound "more alive": overdoing it makes it sound less natural.
- Skipping captions on social video.
Sources
- YouTube Help: Disclosing altered or synthetic content
- YouTube Help: YouTube channel monetization policies
- TikTok Support: AI-generated content
- National Center for Voice and Speech: Voice qualities and speaking rate
- Spotify for Artists: Loudness normalization
- Apple Podcasts for Creators: Audio requirements
- Microsoft Support: Add or delete audio in your PowerPoint presentation
- Microsoft Support: Turn your presentation into a video
- Adobe Help: Automatically duck audio in Premiere
- Larry Jordan: Automatically duck background music under dialog in DaVinci Resolve
- CapCut Help: How to recognise subtitles
- DIY Video Studio: Ideal audio levels for video
- Merriam-Webster: voice-over