Nano GPT logo
NanoGPT
Nano GPT logo
NanoGPT
Nano GPT logo
NanoGPT

Recent

 
 
 
 
All conversations

Create

Chat
Studio

Explore

Models
Pricing
AI Apps
Blog

Library

Gallery
Workspace
Usage
API
Suggestions and bugs
Terms

|

Privacy

|

Refunds
Balance
Subscription
Teams
Invitations
Presets
Settings
FAQPartnersPress & mediaBrand assetsBenchmarksEarn
  1. Models
  2. Audio
  3. Xai

Grok Speech-to-Text

Try Model

SpaceXAI speech-to-text on fal with diarization, word-level timestamps, and multichannel audio support

Details

Approx. Price
$0.003 per minute
Model type
speech-to-text
Category
speech-to-text
Preview examples
0

Related audio models

Compare Grok Speech-to-Text with similar models from the same provider or model family.

Qwen Audio 3.0 TTS Flash

alibaba/qwen-audio-3-tts

Fast multilingual text-to-speech with a broad set of voices and explicit language control.

ElevenLabs Scribe V2

Elevenlabs-Scribe-V2

ElevenLabs Scribe V2 transcription with improved accuracy, word-level timestamps, and speaker identification

ElevenLabs Scribe V1

Elevenlabs-STT

ElevenLabs Scribe V1 transcription with word-level timestamps and speaker identification

ElevenLabs Turbo V2.5

Elevenlabs-Turbo-V2.5

High quality with lowest latency, ideal for real-time applications. Supports 32 languages while maintaining natural voice quality.

ElevenLabs v3

Elevenlabs-V3

High-quality text-to-speech with enhanced controls and natural voices.

Demucs Stem Separation

fal-ai/demucs

Separate up to 60 minutes / 700 MiB of stereo audio into vocals, drums, bass, guitar, piano, and other stems with Demucs.