Fast multilingual text-to-speech with a broad set of voices and explicit language control.
Approx. Price
$85.00 per 1M characters
Model Type
text-to-speech
Category
text-to-speech
Preview Examples
0
Available voices
Built-in voices accepted by NanoGPT for this model. Custom or cloned voice IDs may also be supported.
Related audio models
Compare Qwen Audio 3.0 TTS Flash with similar models from the same provider or model family.
Qwen 3 TTS 1.7B
Qwen-3-TTS-1.7BBring speech to your texts using Qwen3-TTS with pre-trained voices or cloned voice embeddings.
Stable Audio 3 Medium
fal-ai/stable-audio-3/medium/text-to-audioStable Audio 3 Medium generates high-quality stereo music up to 6 minutes from text prompts, trained on fully licensed data for commercial use.
Stable Audio 3 Small Music
fal-ai/stable-audio-3/small/music/text-to-audioStable Audio 3 Small Music generates full stereo music compositions up to 2 minutes from text prompts.
Stable Audio 3 Small SFX
fal-ai/stable-audio-3/small/sfx/text-to-audioStable Audio 3 Small SFX generates high-quality sound effects from text prompts, with controllable clip length, output format, negative prompt, and seed.
ElevenLabs Scribe V2
Elevenlabs-Scribe-V2ElevenLabs Scribe V2 transcription with improved accuracy, word-level timestamps, and speaker identification
ElevenLabs Scribe V1
Elevenlabs-STTElevenLabs Scribe V1 transcription with word-level timestamps and speaker identification