Alibaba Cloud DashScope non-realtime speech recognition with multilingual transcription, punctuation, and sentence/word timestamps.
Approx. Price
$0.004 per minute
Model Type
speech-to-text
Category
speech-to-text
Preview Examples
0
Related audio models
Compare Alibaba Fun-ASR Flash with similar models from the same provider or model family.
Qwen Audio 3.0 TTS Flash
alibaba/qwen-audio-3-ttsFast multilingual text-to-speech with a broad set of voices and explicit language control.
Gemini 2.5 Flash Preview TTS
gemini-2.5-flash-preview-ttsGoogle Gemini native TTS. Single and multi-speaker support via prompt.
Gemini 3.1 Flash TTS Preview
gemini-3.1-flash-tts-previewGoogle Gemini 3.1 Flash text-to-speech with inline audio tag and multi-speaker prompt support.
NVIDIA Nemotron ASR Multilingual
nvidia/nemotron-asr-multilingual/asrOpen multilingual streaming speech recognition model with low latency, native punctuation, and capitalization.
ACE-Step 1.5
ACE-Step-1.5ACE-Step 1.5 composes complete songs from text descriptions. Guide genre, mood, and structure with style tags and required custom lyrics. Generates up to 4 minutes of multi-track audio with vocals.
ACE-Step v1.5 Base
ACE-Step-v1.5-BaseACE-Step v1.5 Base is a Runware-hosted music model built for creator workflows, with stronger fidelity, more reliable stylistic consistency, and prompt-driven genre control for full-song generation.