Fast multilingual text-to-speech with a broad set of voices and explicit language control.
Approx. Price
$85.00 per 1M characters
Model Type
text-to-speech
Category
text-to-speech
Preview Examples
0
Available voices
Built-in voices accepted by NanoGPT for this model. Custom or cloned voice IDs may also be supported.
Related audio models
Compare Qwen Audio 3.0 TTS Flash with similar models from the same provider or model family.
Gemini 3.8 Flash TTS
google/gemini-3.8-flash-ttsExpressive narration and two-speaker dialogue with 30 voices, delivery instructions, and inline vocal events. Returns WAV audio. Billed by exact spoken characters with no per-request minimum; instructions are not billed. The complete input is limited to 8,192 tokens.
Qwen 3 TTS 1.7B
Qwen-3-TTS-1.7BBring speech to your texts using Qwen3-TTS with pre-trained voices or cloned voice embeddings.
Stable Audio 3 Medium
fal-ai/stable-audio-3/medium/text-to-audioStable Audio 3 Medium generates high-quality stereo music up to 6 minutes from text prompts, trained on fully licensed data for commercial use.
Stable Audio 3 Small Music
fal-ai/stable-audio-3/small/music/text-to-audioStable Audio 3 Small Music generates full stereo music compositions up to 2 minutes from text prompts.
Stable Audio 3 Small SFX
fal-ai/stable-audio-3/small/sfx/text-to-audioStable Audio 3 Small SFX generates high-quality sound effects from text prompts, with controllable clip length, output format, negative prompt, and seed.
ElevenLabs Scribe V2
Elevenlabs-Scribe-V2ElevenLabs Scribe V2 transcription with improved accuracy, word-level timestamps, and speaker identification