Alibaba Cloud DashScope non-realtime speech recognition with multilingual transcription, punctuation, and sentence/word timestamps.
Approx. Price
$0.004 per minute
Model Type
speech-to-text
Category
speech-to-text
Preview Examples
0
Related audio models
Compare Alibaba Fun-ASR Flash with similar models from the same provider or model family.
Qwen Audio 3.0 TTS Flash
alibaba/qwen-audio-3-ttsFast multilingual text-to-speech with a broad set of voices and explicit language control.
Gemini 2.5 Flash Preview TTS
gemini-2.5-flash-preview-ttsGoogle Gemini native TTS. Single and multi-speaker support via prompt.
Gemini 3.1 Flash TTS Preview
gemini-3.1-flash-tts-previewGoogle Gemini 3.1 Flash text-to-speech with inline audio tag and multi-speaker prompt support.
Gemini 3.8 Flash TTS
google/gemini-3.8-flash-ttsExpressive narration and two-speaker dialogue with 30 voices, delivery instructions, and inline vocal events. Returns WAV audio. Billed by exact spoken characters with no per-request minimum; instructions are not billed. The complete input is limited to 8,192 tokens.
NVIDIA Nemotron ASR Multilingual
nvidia/nemotron-asr-multilingual/asrOpen multilingual streaming speech recognition model with low latency, native punctuation, and capitalization.
ACE-Step 1.5
ACE-Step-1.5ACE-Step 1.5 composes complete songs from text descriptions. Guide genre, mood, and structure with style tags and required custom lyrics. Generates up to 4 minutes of multi-track audio with vocals.