How to Choose an AI Voice for Text-to-Speech
A voice that sounds impressive in a two-second greeting can become tiring halfway through a long narration. It may also stumble on the product name, rush a date, flatten a joke, or change character after a few paragraphs.
NanoGPT places a play button beside supported voices in the text-to-speech picker, so you can hear a short sample before generating a full clip. That makes the first comparison faster, but a short preview should only be the start of the audition.
Open text-to-speech on NanoGPT.
Match the voice to the job
The same voice can be pleasant in a 15-second announcement and distracting in a 30-minute article. Decide what the listener will hear before comparing names and labels.
- Long narration should be checked for consistent volume, suitable expression, clear pauses, and low listener fatigue. Fiction may need more range than an instructional article.
- Assistants and support benefit from clarity and a delivery that fits the audience and brand without sounding theatrical or impatient.
- Characters and dialogue may need distinct voices, emotional range, and enough separation that listeners can follow who is speaking.
- Ads and short clips can use more energy, but brand names, prices, and calls to action still need precise pronunciation.
- Multilingual speech requires testing every target language and any code-switching that will appear in the final script.
Age, gender, and style labels can help narrow a long list, but they do not tell you whether a voice fits the actual text. Listen before attaching a role to it.
What a short preview tells you
The preview button is useful for checking basic tone, accent, pitch, and general delivery. These are individual built-in previews from several voice-and-model combinations, not a controlled comparison or a demonstration of each library's full range:
- Play Alloy preview — MP3, 1.8 seconds, “Hi, I'm Alloy.”
- Play Aria preview — MP3, 1.4 seconds, “Hi, I'm Aria.”
- Play Rex preview — MP3, 1.3 seconds, “Hi, I'm Rex.”
- Play Sweety Spanish preview — MP3, 1.7 seconds, “Hola, soy Sweety.”
- Play Sohee preview — MP3, 1.8 seconds, “Hi, I'm Sohee.”
These clips can remove obvious mismatches. They cannot test long pauses, difficult names, numbers, sustained emotion, or consistency across several minutes.
Use the same audition script
Once you have a shortlist, generate the same passage with each voice. When comparing voices within one model, keep shared settings fixed. When comparing different models, start with their defaults, then tune each finalist for the intended use. The practical comparison is always a voice, model, and settings combination rather than a voice in isolation.
An audition script should include the material most likely to expose a problem. For a general-purpose English voice, this is a useful starting point:
Welcome back. Your order number is A-204, and it should arrive by Friday, 17 July. Before we continue, here's one detail people sometimes miss: the smaller plan includes offline access. Ready? Let's take it from the top.
That passage checks letters and numbers, a date, two sentence lengths, a colon, a question, contractions, and a change from informative to conversational delivery. Replace its details with your hardest terms: product names, prices, URLs, acronyms, foreign words, and any numbers that must be read exactly.
Decide the expected reading before scoring ambiguous text. For example, A-204 might need to be read as “A two oh four” rather than “A dash two hundred and four.”
For multilingual work, include complete sentences in every target language and passages that switch between languages where your script does. Ask a fluent reviewer who knows the target locale to check pronunciation and rhythm. A multilingual label does not establish equal quality across every listed language.
Listen for specific failures
Comparing voices by whether they sound “good” makes the decision vague. Listen for observable problems instead:
- Are names, dates, prices, and abbreviations understandable on the first listen?
- Do commas, full stops, and paragraph breaks create useful pauses?
- Does the voice rush short words or stretch the final word of a sentence?
- Are sharp consonants uncomfortable on headphones?
- Does the voice stay consistent when the passage becomes longer or more emotional?
- Does it remain pleasant at the volume and on the device your audience will use?
For narration, test at least a few minutes. Listener fatigue rarely appears in a greeting. For an assistant, test several unrelated replies so every answer does not sound like the same recorded announcement. For a character, test calm, excited, uncertain, and interrupted lines rather than one dramatic monologue.
Test the voice and model together
A familiar voice label may appear across model versions or libraries, but the same label does not guarantee the same sound, pronunciation, controls, or response speed.
Choose the model based on the job as well as the available voices. Different models expose different controls. Open the settings panel to see whether speed, language, style instructions, stability, similarity, expressiveness, multi-speaker dialogue, or custom voice IDs are available. Some models expose only a voice and text box.
The meaning and scale of these controls also vary by model. Change one control at a time during the audition. Increasing speed can make a calm voice sound hurried. Stronger style settings can improve a short character line while making long narration uneven. A voice that seems wrong at the default settings may only need a small pacing adjustment.
Before committing to a model for ongoing work, also check generation cost, response speed for interactive uses, language availability, script-length limits, commercial-use terms, privacy rules, and whether the voice is likely to remain available.
A repeatable audition
Open the voice list and use the play buttons to remove obvious mismatches. Keep three candidates, then generate 30 to 60 seconds of the same real script with each one. Use that short round to reduce the shortlist, not to judge long-term comfort.
For the finalists, render several representative minutes or the hardest complete section. Listen at roughly the same loudness and on the device your audience is likely to use. Score clarity, pacing, fit, pronunciation, and fatigue separately on a simple 1-to-5 scale instead of relying on one overall impression. If several people are reviewing the samples, shuffle them, rename the files to neutral labels such as sample_a.mp3, and reveal the voice names after scoring.
Run the longest or most difficult section before making the final choice. A voice selected for a series, course, game, or assistant can be costly and time-consuming to replace after listeners associate it with the product.
Custom and cloned voices
Where supported, NanoGPT text-to-speech models can accept saved custom voice IDs or cloned-voice inputs. These can create a recognizable identity, although consistency still needs testing. A good reference recording does not guarantee every generated line will preserve pacing, emotion, or pronunciation.
Before cloning or deliberately imitating an identifiable person, obtain documented authorization or another lawful basis for the intended use, confirm that the model's terms permit it, and protect the reference recording. Disclose synthetic speech where law or platform rules require it and whenever listeners could reasonably mistake it for a real recording.