Skip to navigation

Create transcription

Transcribes a recording and answers with two transcripts: verbatim, what was said, including fillers, repetitions and cut-offs; and intended, what was meant, as fluent text. Both carry word-level timestamps, and alignment relates the two word by word. Send language when you know it. The quickstart walks through the request, the accepted formats, the supported languages and every field of the response.

Authentication

AuthorizationBearer

Your API key, e.g. Authorization: Bearer nl_live_....

Request

This endpoint expects a multipart form containing a file.
filefileRequired

The recording to transcribe, up to 25 MB. Accepted formats, detected from the bytes rather than the filename: wav, mp3, m4a (AAC in an MP4 container), flac, ogg (Vorbis or Opus) and webm. Video files are refused; send the audio track.

languageenumOptional

The language of the audio as an ISO-639-1 code, e.g. en or de. Optional but recommended: when it is omitted the language is detected from the first 30 seconds, which costs a little time and can guess wrong on short or noisy recordings. See the supported languages: German, English, Spanish, French, Portuguese, Russian, Swedish, Dutch, Polish, Italian, Danish and Ukrainian are officially supported; the other codes are available in beta. A code outside the list answers unsupported_language.

modelenumRequired

The model to transcribe with. Required; crisperwhisper-v2 is the only model, and any other value answers model_not_found.

Allowed values:

Response

The transcription.
alignmentlist of objects or null

The verbatim transcript aligned against the intended one, in reading order. null when the two could not be aligned; the transcripts are still valid.

durationdouble
Duration of the decoded audio, in seconds.
intendedobject

What the speaker meant: fluent, without fillers, repetitions or cut-offs.

languagestring

ISO-639-1 code of the audio's language, as sent or as detected. See the supported languages: twelve are officially supported, the rest are available in beta.

usageobject
verbatimobject

What the speaker said, including fillers, repetitions and cut-offs.

Errors

400
Bad Request Error
401
Unauthorized Error
403
Forbidden Error
413
Content Too Large Error
429
Too Many Requests Error
503
Service Unavailable Error