Skip to navigation

Quickstart

One request, two transcripts: what was said and what was meant, with word timestamps.

POST /v1/audio/transcriptions takes an audio file and returns a verbatim transcript (fillers, repetitions and cut-offs included), an intended transcript (the fluent version), word timestamps for both, and an alignment that links the two.

1

Get a key

Create one in the dashboard. The plaintext is shown once.

export NYRA_API_KEY="nl_live_YxQFoF6p2xKvStMzWhYhwPMNSlEnWTwzYL83OJCbyTo"
2

Top up the wallet

Billing is prepaid. A key with an empty wallet answers 429 insufficient_quota. See Billing.

3

Transcribe

Send the file as multipart/form-data.

curl https://api.nyra-labs.com/v1/audio/transcriptions \
-H "Authorization: Bearer $NYRA_API_KEY" \
-F "file=@clip.wav" \
-F "model=crisperwhisper-v2" \
-F "language=en"

Parameters

The request is multipart/form-data; any other body answers 400 invalid_content_type before it is read.

FieldNotes
filerequiredOne audio file, up to 25 MB, in one of the accepted formats. Exactly one file part; a text field, a missing part or a second file is a 400.
modelrequiredcrisperwhisper-v2. Any other value answers 400 model_not_found; GET /v1/models lists what is available.
languageoptional, recommendedISO-639-1 code from the supported list, e.g. en, de. Anything else answers 400 unsupported_language.

Accepted formats

  • wav
  • mp3
  • m4a (AAC)
  • flac
  • ogg (Vorbis or Opus)
  • webm (audio only)

Any other format, and any file with a video stream, answers 400 unsupported_file_type. The model works on 16 kHz mono audio, so to keep requests no larger than necessary, resample your audio to 16 kHz before you send it.

Supported languages

language is an ISO-639-1 code (three letters where no two-letter code exists). These languages are officially supported:

CodeLanguage
deGerman
enEnglish
esSpanish
frFrench
ptPortuguese
ruRussian
svSwedish
nlDutch
plPolish
itItalian
daDanish
ukUkrainian

The following languages are available in beta:

CodeLanguage
afAfrikaans
sqAlbanian
amAmharic
arArabic
hyArmenian
asAssamese
azAzerbaijani
baBashkir
euBasque
beBelarusian
bnBengali
bsBosnian
brBreton
bgBulgarian
myBurmese
yueCantonese
caCatalan
zhChinese
hrCroatian
csCzech
etEstonian
foFaroese
fiFinnish
glGalician
kaGeorgian
elGreek
guGujarati
htHaitian Creole
haHausa
hawHawaiian
heHebrew
hiHindi
huHungarian
isIcelandic
idIndonesian
jaJapanese
jwJavanese
knKannada
kkKazakh
kmKhmer
koKorean
loLao
laLatin
lvLatvian
lnLingala
ltLithuanian
lbLuxembourgish
mkMacedonian
mgMalagasy
msMalay
mlMalayalam
mtMaltese
miMaori
mrMarathi
mnMongolian
neNepali
noNorwegian
nnNorwegian Nynorsk
ocOccitan
psPashto
faPersian
paPunjabi
roRomanian
saSanskrit
srSerbian
snShona
sdSindhi
siSinhala
skSlovak
slSlovenian
soSomali
suSundanese
swSwahili
tlTagalog
tgTajik
taTamil
ttTatar
teTelugu
thThai
boTibetan
trTurkish
tkTurkmen
urUrdu
uzUzbek
viVietnamese
cyWelsh
yiYiddish
yoYoruba

Case does not matter (DE is de). Region subtags are not accepted: send pt, not pt-BR.

Send language when you know it. When it is omitted, the model spends an extra decoding pass on the first 30 seconds of audio to detect the language, and a wrong guess (a short clip, a noisy opening, a name in another language) degrades the whole transcript. When you pass it, that pass is skipped and the result is the language you sent.

The response

For a clip in which someone says “I I want um to go home”:

{
"language": "en",
"duration": 2.84,
"verbatim": {
"text": "I I want um to go home",
"words": [
{ "word": " I", "start": 0.12, "end": 0.26, "alignment_index": 0 },
{ "word": " I", "start": 0.34, "end": 0.48, "alignment_index": 1 },
{ "word": " want", "start": 0.52, "end": 0.81, "alignment_index": 2 },
{ "word": " um", "start": 0.95, "end": 1.31, "alignment_index": 3 },
{ "word": " to", "start": 1.58, "end": 1.69, "alignment_index": 4 },
{ "word": " go", "start": 1.72, "end": 1.94, "alignment_index": 5 },
{ "word": " home", "start": 1.98, "end": 2.41, "alignment_index": 6 }
]
},
"intended": {
"text": "I want to go home",
"words": [
{ "word": " I", "start": 0.34, "end": 0.48, "alignment_index": 1 },
{ "word": " want", "start": 0.52, "end": 0.81, "alignment_index": 2 },
{ "word": " to", "start": 1.58, "end": 1.69, "alignment_index": 4 },
{ "word": " go", "start": 1.72, "end": 1.94, "alignment_index": 5 },
{ "word": " home", "start": 1.98, "end": 2.41, "alignment_index": 6 }
]
},
"alignment": [
{ "type": "repetition", "verbatim": "I", "intended": null },
{ "type": "match", "verbatim": "I", "intended": "I" },
{ "type": "match", "verbatim": "want", "intended": "want" },
{ "type": "filler", "verbatim": "um", "intended": null },
{ "type": "match", "verbatim": "to", "intended": "to" },
{ "type": "match", "verbatim": "go", "intended": "go" },
{ "type": "match", "verbatim": "home", "intended": "home" }
],
"usage": { "type": "duration", "seconds": 3 }
}

verbatim and intended

Both are a text string plus a words list, and both come from one decode: CrisperWhisper 2 produces the two transcripts together rather than transcribing once and cleaning up afterwards, so they agree on what was in the audio and differ only in what they keep.

  • verbatim is what was said. It keeps fillers (um, uh), repetitions, false starts, and cut-off words. Use it when the manner of speaking is the point: clinical assessment, speech therapy, fluency research, subtitles that must be faithful.
  • intended is what was meant: the same speech read as fluent text. Use it for notes, search, summaries, or anything downstream that would trip over a stutter.

Words and timestamps

Each word has a start and end in seconds from the start of the recording, and keeps its leading space. Concatenating the words reproduces the track’s text:

assert "".join(w["word"] for w in track["words"]).strip() == track["text"]

The leading space is how a word tells you whether it began a new token or continued one; keep it if you rebuild text, strip it if you display words on their own.

Timestamps are per track: each transcript is timed from its own decode, so a word that appears in both tracks (the second I above) normally carries the same timing in both, but the two are not guaranteed to be identical to the millisecond.

alignment and alignment_index

alignment is the link between the two tracks, in reading order. Each entry names the token on the verbatim side, the token on the intended side (or null if there is none), and how they relate:

typeMeaning
matchThe same word on both sides.
repetitionA repeated word or phrase; only the final instance is kept in intended.
disfluentA false start or restart that does not appear in intended.
cutoffA word the speaker broke off mid-way.
fillerum, uh, er and the like.
soundA non-word vocalisation: a laugh, a cough, a click.
correctionThe speaker corrected themselves; verbatim is what was first said, intended is what replaced it.

Each word’s alignment_index is the index of its entry in alignment, so you can walk from a verbatim word to its intended counterpart (or learn that it has none) without string matching. It is null when the classifier could not prove the link. A null is never a guess: the API does not invent links it cannot show.

alignment itself is null when the classifier could not run on this clip, which happens on very long or very dense recordings that exceed its bounds. That is not an error: verbatim and intended are complete and correct without it, and the request is billed normally.

Audio without speech

Silence, music or noise is a valid request, not an error: both transcripts come back with an empty text and an empty words list, alignment is [], and the audio is billed like any other. Check verbatim.text before assuming something was said.

usage

usage.seconds is the audio’s decoded duration rounded up to a whole second: what you were charged for, at 0.90peraudiohour(0.90 per audio hour (0.00025 per second). The 2.84 s clip above is 3 seconds, $0.00075. See Billing.

Python

import os
import requests
with open("clip.wav", "rb") as audio:
response = requests.post(
"https://api.nyra-labs.com/v1/audio/transcriptions",
headers={"Authorization": f"Bearer {os.environ['NYRA_API_KEY']}"},
files={"file": ("clip.wav", audio, "audio/wav")},
data={"model": "crisperwhisper-v2", "language": "en"},
timeout=600,
)
response.raise_for_status()
result = response.json()
print(result["verbatim"]["text"])
print(result["intended"]["text"])
for word in result["verbatim"]["words"]:
print(f"{word['start']:6.2f} {word['end']:6.2f} {word['word']}")

Node

import { openAsBlob } from "node:fs";
const form = new FormData();
form.append("file", await openAsBlob("clip.wav"), "clip.wav");
form.append("model", "crisperwhisper-v2");
form.append("language", "en");
const response = await fetch("https://api.nyra-labs.com/v1/audio/transcriptions", {
method: "POST",
headers: { Authorization: `Bearer ${process.env.NYRA_API_KEY}` },
body: form,
});
if (!response.ok) {
throw new Error(`${response.status}: ${await response.text()}`);
}
const result = await response.json();
console.log(result.verbatim.text);
console.log(result.intended.text);

Limits

LimitOver it
Upload size25 MB413 file_too_large
Audio length1 hour of decoded audio400 audio_too_long
Formatswav, mp3, m4a, flac, ogg, webm (details)400 unsupported_file_type (readable audio in another format, or a video); 400 invalid_audio (not audio at all)
Requests600 per minute per key by default; adjustable per key in the dashboard429 rate_limit_exceeded

Both the size and the length are checked before any money is reserved and before the audio reaches a model. Split long recordings client-side; a 25 MB budget fits about 13 minutes of 16 kHz mono WAV or well over an hour of compressed audio, so which limit you meet first depends on the format.

Response headers

HeaderMeaning
x-request-idThe req_... id. Quote it in support requests; the dashboard’s request explorer is keyed by it.
x-nyra-audio-duration-msThe decoded duration, in integer milliseconds.
x-nyra-cost-usdWhat this request cost, in USD as a fixed-point decimal.

Rate-limit headers (x-ratelimit-limit-requests, x-ratelimit-remaining-requests, x-ratelimit-reset-requests) are on every response too; see Authentication.

Retries and idempotency

The endpoint accepts an Idempotency-Key header of up to 255 characters. Once a request with a given key has completed, a retry with the same key within 24 hours replays the stored response: the same body and the same x-request-id, plus x-nyra-idempotent-replay: true, and nothing is charged again.

curl https://api.nyra-labs.com/v1/audio/transcriptions \
-H "Authorization: Bearer $NYRA_API_KEY" \
-H "Idempotency-Key: 6f1c0b4e-2f19-4f0e-b2f1-9a3c5d7e8f01" \
-F "file=@clip.wav" -F "model=crisperwhisper-v2" -F "language=en"

Two things it is not:

  • It is not a lock. Two genuinely concurrent requests with the same key both run and both charge. Wait for the first to finish before retrying.
  • It is not a cache under zero retention. If your organization stores no transcripts, the response is not kept either; a retry answers 400 idempotency_response_expired without running or charging again.

A key is bound to the model and language of its first request: a retry under the same key with a different language or model answers 400 idempotency_key_in_use rather than the other request’s transcript.

Keys are namespaced to your organization, so nobody else’s key collides with yours.

Errors

Every non-2xx response carries the same envelope:

{
"error": {
"message": "The audio is longer than the 3600 second limit. Split it into shorter files.",
"type": "invalid_request_error",
"param": "file",
"code": "audio_too_long"
}
}

code is the field to branch on; message is for humans and may change. param names the offending field when there is one.

StatuscodeWhen
400invalid_content_typeThe body is not multipart/form-data.
400missing_required_parameterfile or model was not sent; param names it.
400invalid_valuefile is a text field rather than an upload, or more than one file was sent.
400invalid_audioThe file is empty, or its bytes are not audio in any container we can read.
400unsupported_file_typeReadable audio outside the accepted formats, or a file with a video stream.
400audio_too_longThe decoded audio is longer than one hour.
400unsupported_languagelanguage is not in the supported list.
400unsupported_parameterA removed field such as response_format was sent; param names it.
400model_not_foundmodel names a model that does not exist or is retired.
400idempotency_response_expiredSee Retries and idempotency.
400idempotency_key_in_useThe same Idempotency-Key was used with a different language or model.
401missing_api_key, invalid_api_keyNo key, or a key that is unknown, revoked or expired.
403insufficient_scopeThe key was narrowed and lacks transcriptions:write.
413file_too_largeThe upload is over 25 MB.
429insufficient_quotaThe wallet cannot cover the request’s hold. Retrying will not help; top up.
429rate_limit_exceededToo many requests in the current minute. retry-after says how long to wait.
503capacity_exceededThe transcription backend is saturated. Nothing was charged. Retry after the retry-after header says.

insufficient_quota and rate_limit_exceeded share a status code but not a remedy: back off on the second, stop and top up on the first. Check code, not the status.

The API reference is generated from the same OpenAPI spec the server is built from. If a behaviour is not in the spec, it is not part of the public API.