Skip to main content
POST
Speech to Text

Request

Send as multipart/form-data.

Form parameters

file
required
Audio file to transcribe. Supported formats: mp3, mp4, mpeg, mpga, m4a, wav, webm. Max size: 25 MB.
string
STT model ID. Options: whisper-large-v3, whisper-large-v3-turbo, sarvam-stt. Default: whisper-large-v3-turbo.
string
Language code (ISO 639-1). Example: "en", "hi", "es".
string
Output format. Options: "json", "text", "verbose_json". Default: "json".
number
Sampling temperature for the model. Default: 0.
string
Optional prompt to guide the transcription style.

Examples


Response

Pricing

All STT models are premium and require wallet balance.