Recent
Create
Explore
Library
|
SpaceXAI speech-to-text on fal with diarization, word-level timestamps, and multichannel audio support
Compare Grok Speech-to-Text with similar models from the same provider or model family.
Qwen Audio 3.0 TTS Flash
alibaba/qwen-audio-3-tts
Fast multilingual text-to-speech with a broad set of voices and explicit language control.
ElevenLabs Scribe V2
Elevenlabs-Scribe-V2
ElevenLabs Scribe V2 transcription with improved accuracy, word-level timestamps, and speaker identification
ElevenLabs Scribe V1
Elevenlabs-STT
ElevenLabs Scribe V1 transcription with word-level timestamps and speaker identification
ElevenLabs Turbo V2.5
Elevenlabs-Turbo-V2.5
High quality with lowest latency, ideal for real-time applications. Supports 32 languages while maintaining natural voice quality.
ElevenLabs v3
Elevenlabs-V3
High-quality text-to-speech with enhanced controls and natural voices.
Demucs Stem Separation
fal-ai/demucs
Separate up to 60 minutes / 700 MiB of stereo audio into vocals, drums, bass, guitar, piano, and other stems with Demucs.