Create transcription
Transcribes a recording and answers with two transcripts: verbatim, what was said, including fillers, repetitions and cut-offs; and intended, what was meant, as fluent text. Both carry word-level timestamps, and alignment relates the two word by word. Send language when you know it. The quickstart walks through the request, the accepted formats, the supported languages and every field of the response.
Authentication
Your API key, e.g. Authorization: Bearer nl_live_....
Request
The recording to transcribe, up to 25 MB. Accepted formats, detected from the bytes rather than the filename: wav, mp3, m4a (AAC in an MP4 container), flac, ogg (Vorbis or Opus) and webm. Video files are refused; send the audio track.
The language of the audio as an ISO-639-1 code, e.g. en or de. Optional but recommended: when it is omitted the language is detected from the first 30 seconds, which costs a little time and can guess wrong on short or noisy recordings. See the supported languages: German, English, Spanish, French, Portuguese, Russian, Swedish, Dutch, Polish, Italian, Danish and Ukrainian are officially supported; the other codes are available in beta. A code outside the list answers unsupported_language.
The model to transcribe with. Required; crisperwhisper-v2 is the only model, and any other value answers model_not_found.
Response
The verbatim transcript aligned against the intended one, in reading order. null when the two could not be aligned; the transcripts are still valid.
What the speaker meant: fluent, without fillers, repetitions or cut-offs.
ISO-639-1 code of the audio's language, as sent or as detected. See the supported languages: twelve are officially supported, the rest are available in beta.
What the speaker said, including fillers, repetitions and cut-offs.