Microsoft high-fidelity expressive text-to-speech with multilingual MAI-Voice-2 prebuilt voices.
Approx. Price
$37.40 per 1M characters
Model Type
text-to-speech
Category
text-to-speech
Preview Examples
0
Related audio models
Compare MAI-Voice-2 with similar models from the same provider or model family.
MAI-Transcribe 1.5
microsoft/mai-transcribe-1.5Microsoft fast transcription model with automatic language detection, punctuation, and 100+ BCP-47 locales.
VibeVoice
microsoft/vibevoiceLong-form text-to-speech with multi-speaker dialogue support and 9 voice presets across English, Chinese, and Hindi.
ACE-Step 1.5
ACE-Step-1.5ACE-Step 1.5 composes complete songs from text descriptions. Guide genre, mood, and structure with style tags and required custom lyrics. Generates up to 4 minutes of multi-track audio with vocals.
ACE-Step v1.5 Base
ACE-Step-v1.5-BaseACE-Step v1.5 Base is a Runware-hosted music model built for creator workflows, with stronger fidelity, more reliable stylistic consistency, and prompt-driven genre control for full-song generation.
ACE-Step v1.5 Turbo
ACE-Step-v1.5-TurboACE-Step v1.5 Turbo is the faster, lower-cost Runware variant for full-song generation, with broad genre coverage, improved stylistic consistency, and text-guided music creation for creator workflows.
Qwen Audio 3.0 TTS Flash
alibaba/qwen-audio-3-ttsFast multilingual text-to-speech with a broad set of voices and explicit language control.