Fast multilingual text-to-speech with a broad set of voices and explicit language control.
Approx. Price
$85.00 per 1M characters
Model Type
text-to-speech
Category
text-to-speech
Preview Examples
0
Related audio models
Compare Qwen Audio 3.0 TTS Flash with similar models from the same provider or model family.
Qwen 3 TTS 1.7B
Qwen-3-TTS-1.7BBring speech to your texts using Qwen3-TTS with pre-trained voices or cloned voice embeddings.
Stable Audio 3 Medium
fal-ai/stable-audio-3/medium/text-to-audioStable Audio 3 Medium generates high-quality stereo music up to 6 minutes from text prompts, trained on fully licensed data for commercial use.
Stable Audio 3 Small Music
fal-ai/stable-audio-3/small/music/text-to-audioStable Audio 3 Small Music generates full stereo music compositions up to 2 minutes from text prompts.
Stable Audio 3 Small SFX
fal-ai/stable-audio-3/small/sfx/text-to-audioStable Audio 3 Small SFX generates high-quality sound effects from text prompts, with controllable clip length, output format, negative prompt, and seed.
ElevenLabs Scribe V2
Elevenlabs-Scribe-V2ElevenLabs Scribe V2 transcription with improved accuracy, word-level timestamps, and speaker identification
ElevenLabs Scribe V1
Elevenlabs-STTElevenLabs Scribe V1 transcription with word-level timestamps and speaker identification