Seedance 2.5 is here. Every model is included on every plan.
Dolly
Sign inSign up

How to choose an AI voice (without listening to 300 of them)

5 min read

A fast method for picking the right AI voice: filter by gender, age, accent and use case, preview a shortlist, then test with your own script before committing.

The hard part of AI voiceover is no longer quality. Current models read naturally enough that listeners rarely notice. The hard part is choice: with hundreds of voices available, "just pick one" turns into forty minutes of aimless previewing, and the voice you settle on is whichever one you heard last.

Here is a faster method. It takes about five minutes and ends with a voice you can defend.

Decide three things before you listen

Write down, in one line, who should be speaking. Not a name, a casting note: "middle aged American woman, warm, sounds like she has done this before" or "young British man, quick, a bit dry." The three attributes doing the work are age, accent and energy. If you cannot fill in that sentence, the previews will not decide it for you; look at three videos you admire in your niche and describe their narrators.

Filter first, listen second

In the voice studio, the picker holds around 300 voices, nearly all carrying real metadata: gender, age, accent, language and intended use case, plus a short character description. Apply your casting note as filters. "Female, old" leaves a few dozen candidates; add "American" and you are looking at nine voices, among them a likeable grandma, a grandmotherly storykeeper and a little old lady you would never find by scrolling. The filters exist because this data was captured per voice, so what you select on is what the voice actually is.

Now preview, but only within the filtered set, and only long enough to sort each voice into yes or no. Every voice has a free sample a click away. Your shortlist should be three to five voices.

Test with your own script

Preview samples are performed in each voice's comfort zone. Your script is not that. Run one real paragraph of your actual script through each shortlisted voice. This costs a dime or two per test (voice generation is priced per started 1,000 characters of script, 8 to 15 credits on the ElevenLabs tiers), and it is where the decision actually happens. Listen for three failure points:

  • Product names and jargon. Odd words are where synthetic voices stumble first. If your product name comes out wrong, respelling it phonetically in the script usually fixes it.
  • Sentence endings. Weaker matches develop a repetitive falling cadence that gets noticeable over a full minute.
  • Numbers and lists. Pricing, steps, specs. If your content is heavy on these, test them specifically.

Write for the ear

Whichever voice wins, scripts read aloud obey different rules than text. Short sentences. Contractions everywhere. One idea per sentence, and a real pause where a listener needs to catch up (a period does this; a comma does not). On expressive models you can go further: the V3 tier accepts inline tags like [chuckles] or [slowly], which render as actual performance rather than spoken words. The voiceover guide covers script craft in more depth, including multi-speaker dialogue.

Keep a house voice

Once a voice works, reuse it. A consistent narrator across your videos does the same job a logo does, and swapping voices between videos in a series reads as sloppy. Note the voice name somewhere your team can find it. When you eventually need a second voice (a different product line, a different audience), run this method again rather than drifting to whatever sounds fresh that day.