Private AI
Private AI
Browse and discover the best AI audio models for text to speech, speech to text, and music.
ACE-Step composes complete songs from text descriptions. Guide genre, mood, and structure with style tags and optional custom lyrics. Generates up to 4 minutes of multi-track audio with vocals.
≈ $0.050 per audio
ACE-Step 1.5 composes complete songs from text descriptions. Guide genre, mood, and structure with style tags and required custom lyrics. Generates up to 4 minutes of multi-track audio with vocals.
≈ $0.050 per audio
ACE-Step v1.5 Base is a Runware-hosted music model built for creator workflows, with stronger fidelity, more reliable stylistic consistency, and prompt-driven genre control for full-song generation.
≈ $0.009 per audio
ACE-Step v1.5 Turbo is the faster, lower-cost Runware variant for full-song generation, with broad genre coverage, improved stylistic consistency, and text-guided music creation for creator workflows.
≈ $0.006 per audio
Alibaba Cloud DashScope non-realtime speech recognition with multilingual transcription, punctuation, and sentence/word timestamps.
≈ $0.004 per minute
ByteDance Seed Audio 1.0 generates natural audio from text, with optional preset voices, up to three reference audio clips, or a single reference image.
≈ $300.06 per 1M characters (billed in 300-character blocks)
ByteDance Seed Speech TTS 2.0 for natural multilingual speech with voice instructions and delivery controls.
≈ $51.00 per 1M characters (billed in 1k-character blocks)
Separate up to 60 minutes / 700 MiB of stereo audio into vocals, drums, bass, guitar, piano, and other stems with Demucs.
≈ $0.023 per audio
ElevenLabs Scribe V1 transcription with word-level timestamps and speaker identification
≈ $0.051 per minute
ElevenLabs Scribe V2 transcription with improved accuracy, word-level timestamps, and speaker identification
≈ $0.051 per minute
Generate sound effects and seamless loops from a text description.
≈ $0.010 per audio
High quality with lowest latency, ideal for real-time applications. Supports 32 languages while maintaining natural voice quality.
≈ $102.00 per 1M characters
High-quality text-to-speech with enhanced controls and natural voices.
≈ $170.00 per 1M characters
Google Gemini native TTS. Single and multi-speaker support via prompt.
≈ $51.17 per 1M characters
Higher-quality Gemini TTS with controllable style and tone.
≈ $102.17 per 1M characters
Google Gemini 3.1 Flash text-to-speech with inline audio tag and multi-speaker prompt support.
≈ $102.42 per 1M characters
Google Lyria 3 Pro generates premium music clips from a text prompt, with optional image guidance, negative prompts, and seed-based repeatability.
≈ $0.080 per audio