Long-form text-to-speech with multi-speaker dialogue support and 9 voice presets across English, Chinese, and Hindi.
Approx. Price
$0.150 per generation
Model Type
text-to-speech
Category
text-to-speech
Preview Examples
0
Available voices
Built-in voices accepted by NanoGPT for this model. Custom or cloned voice IDs may also be supported.
Related audio models
Compare VibeVoice with similar models from the same provider or model family.
MAI-Transcribe 1.5
microsoft/mai-transcribe-1.5Microsoft fast transcription model with automatic language detection, punctuation, and 100+ BCP-47 locales.
MAI-Voice-2
microsoft/mai-voice-2Microsoft high-fidelity expressive text-to-speech with multilingual MAI-Voice-2 prebuilt voices.
ACE-Step 1.5
ACE-Step-1.5ACE-Step 1.5 composes complete songs from text descriptions. Guide genre, mood, and structure with style tags and required custom lyrics. Generates up to 4 minutes of multi-track audio with vocals.
ACE-Step v1.5 Base
ACE-Step-v1.5-BaseACE-Step v1.5 Base is a Runware-hosted music model built for creator workflows, with stronger fidelity, more reliable stylistic consistency, and prompt-driven genre control for full-song generation.
ACE-Step v1.5 Turbo
ACE-Step-v1.5-TurboACE-Step v1.5 Turbo is the faster, lower-cost Runware variant for full-song generation, with broad genre coverage, improved stylistic consistency, and text-guided music creation for creator workflows.
Qwen Audio 3.0 TTS Flash
alibaba/qwen-audio-3-ttsFast multilingual text-to-speech with a broad set of voices and explicit language control.