Expressive narration and two-speaker dialogue with 30 voices, delivery instructions, and inline vocal events. Returns WAV audio. The complete input, including instructions, is limited to 8,192 tokens by the provider.
Approx. Price
$76.50 per 1M characters
Model Type
text-to-speech
Category
text-to-speech
Preview Examples
0
Available voices
Built-in voices accepted by NanoGPT for this model. Custom or cloned voice IDs may also be supported.
Related audio models
Compare Gemini 3.8 Flash TTS with similar models from the same provider or model family.
Qwen Audio 3.0 TTS Flash
alibaba/qwen-audio-3-ttsFast multilingual text-to-speech with a broad set of voices and explicit language control.
Qwen 3 TTS 1.7B
Qwen-3-TTS-1.7BBring speech to your texts using Qwen3-TTS with pre-trained voices or cloned voice embeddings.
ElevenLabs Scribe V2
Elevenlabs-Scribe-V2ElevenLabs Scribe V2 transcription with improved accuracy, word-level timestamps, and speaker identification
ElevenLabs Scribe V1
Elevenlabs-STTElevenLabs Scribe V1 transcription with word-level timestamps and speaker identification
ElevenLabs Turbo V2.5
Elevenlabs-Turbo-V2.5High quality with lowest latency, ideal for real-time applications. Supports 32 languages while maintaining natural voice quality.
ElevenLabs v3
Elevenlabs-V3High-quality text-to-speech with enhanced controls and natural voices.