Quickstart
POST /v1/audio/transcriptions takes an audio file and returns a verbatim
transcript (fillers, repetitions and cut-offs included), an intended
transcript (the fluent version), word timestamps for both, and an alignment
that links the two.
Top up the wallet
Billing is prepaid. A key with an empty wallet answers 429 insufficient_quota. See Billing.
Parameters
The request is multipart/form-data; any other body answers
400 invalid_content_type before it is read.
Accepted formats
wavmp3m4a(AAC)flacogg(Vorbis or Opus)webm(audio only)
Any other format, and any file with a video stream, answers
400 unsupported_file_type. The model works on 16 kHz mono audio, so to keep
requests no larger than necessary, resample your audio to 16 kHz before you
send it.
Supported languages
language is an ISO-639-1 code (three letters where no two-letter code
exists). These languages are officially supported:
The following languages are available in beta:
Case does not matter (DE is de). Region subtags are not accepted: send
pt, not pt-BR.
Send language when you know it. When it is omitted, the model spends an
extra decoding pass on the first 30 seconds of audio to detect the language,
and a wrong guess (a short clip, a noisy opening, a name in another language)
degrades the whole transcript. When you pass it, that pass is skipped and the
result is the language you sent.
The response
For a clip in which someone says “I I want um to go home”:
verbatim and intended
Both are a text string plus a words list, and both come from one
decode: CrisperWhisper 2 produces the two transcripts together rather than
transcribing once and cleaning up afterwards, so they agree on what was in the
audio and differ only in what they keep.
verbatimis what was said. It keeps fillers (um, uh), repetitions, false starts, and cut-off words. Use it when the manner of speaking is the point: clinical assessment, speech therapy, fluency research, subtitles that must be faithful.intendedis what was meant: the same speech read as fluent text. Use it for notes, search, summaries, or anything downstream that would trip over a stutter.
Words and timestamps
Each word has a start and end in seconds from the start of the
recording, and keeps its leading space. Concatenating the words
reproduces the track’s text:
The leading space is how a word tells you whether it began a new token or continued one; keep it if you rebuild text, strip it if you display words on their own.
Timestamps are per track: each transcript is timed from its own decode, so a
word that appears in both tracks (the second I above) normally carries the
same timing in both, but the two are not guaranteed to be identical to the
millisecond.
alignment and alignment_index
alignment is the link between the two tracks, in reading order. Each entry
names the token on the verbatim side, the token on the intended side (or
null if there is none), and how they relate:
Each word’s alignment_index is the index of its entry in alignment, so
you can walk from a verbatim word to its intended counterpart (or learn that
it has none) without string matching. It is null when the classifier could
not prove the link. A null is never a guess: the API does not invent links
it cannot show.
alignment itself is null when the classifier could not run on this clip,
which happens on very long or very dense recordings that exceed its bounds.
That is not an error: verbatim and intended are complete and correct
without it, and the request is billed normally.
Audio without speech
Silence, music or noise is a valid request, not an error: both transcripts
come back with an empty text and an empty words list, alignment is [],
and the audio is billed like any other. Check verbatim.text before
assuming something was said.
usage
usage.seconds is the audio’s decoded duration rounded up to a whole
second: what you were charged for, at 0.00025 per
second). The 2.84 s clip above is 3 seconds, $0.00075. See
Billing.
Python
Node
Limits
Both the size and the length are checked before any money is reserved and before the audio reaches a model. Split long recordings client-side; a 25 MB budget fits about 13 minutes of 16 kHz mono WAV or well over an hour of compressed audio, so which limit you meet first depends on the format.
Response headers
Rate-limit headers (x-ratelimit-limit-requests, x-ratelimit-remaining-requests,
x-ratelimit-reset-requests) are on every response too; see
Authentication.
Retries and idempotency
The endpoint accepts an Idempotency-Key header of up to 255 characters.
Once a request with a given key has completed, a retry with the same key
within 24 hours replays the stored response: the same body and the same
x-request-id, plus x-nyra-idempotent-replay: true, and nothing is charged
again.
Two things it is not:
- It is not a lock. Two genuinely concurrent requests with the same key both run and both charge. Wait for the first to finish before retrying.
- It is not a cache under zero retention. If your organization stores no
transcripts, the response is not kept either; a retry answers
400 idempotency_response_expiredwithout running or charging again.
A key is bound to the model and language of its first request: a retry under
the same key with a different language or model answers
400 idempotency_key_in_use rather than the other request’s transcript.
Keys are namespaced to your organization, so nobody else’s key collides with yours.
Errors
Every non-2xx response carries the same envelope:
code is the field to branch on; message is for humans and may change.
param names the offending field when there is one.
insufficient_quota and rate_limit_exceeded share a status code but not a
remedy: back off on the second, stop and top up on the first. Check code,
not the status.
The API reference is generated from the same OpenAPI spec the server is built from. If a behaviour is not in the spec, it is not part of the public API.