> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://docs.nyra-labs.com/quickstart/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nyra-labs.com/_mcp/server. # Quickstart `POST /v1/audio/transcriptions` takes an audio file and returns a **verbatim** transcript (fillers, repetitions and cut-offs included), an **intended** transcript (the fluent version), word timestamps for both, and an alignment that links the two. #### Get a key Create one in the [dashboard](https://platform.nyra-labs.com/api-keys). The plaintext is shown once. ```bash export NYRA_API_KEY="nl_live_YxQFoF6p2xKvStMzWhYhwPMNSlEnWTwzYL83OJCbyTo" ``` #### Top up the wallet Billing is prepaid. A key with an empty wallet answers `429 insufficient_quota`. See [Billing](/billing). #### Transcribe Send the file as `multipart/form-data`. ```bash curl https://api.nyra-labs.com/v1/audio/transcriptions \ -H "Authorization: Bearer $NYRA_API_KEY" \ -F "file=@clip.wav" \ -F "model=crisperwhisper-v2" \ -F "language=en" ``` ## Parameters The request is `multipart/form-data`; any other body answers `400 invalid_content_type` before it is read. | Field | | Notes | | ---------- | --------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `file` | required | One audio file, up to **25 MB**, in one of the [accepted formats](#accepted-formats). Exactly one `file` part; a text field, a missing part or a second file is a `400`. | | `model` | required | `crisperwhisper-v2`. Any other value answers `400 model_not_found`; `GET /v1/models` lists what is available. | | `language` | optional, recommended | ISO-639-1 code from the [supported list](#supported-languages), e.g. `en`, `de`. Anything else answers `400 unsupported_language`. | ### Accepted formats * `wav` * `mp3` * `m4a` (AAC) * `flac` * `ogg` (Vorbis or Opus) * `webm` (audio only) Any other format, and any file with a video stream, answers `400 unsupported_file_type`. The model works on 16 kHz mono audio, so to keep requests no larger than necessary, resample your audio to 16 kHz before you send it. ### Supported languages `language` is an ISO-639-1 code (three letters where no two-letter code exists). These languages are officially supported: | Code | Language | | ---- | ---------- | | `de` | German | | `en` | English | | `es` | Spanish | | `fr` | French | | `pt` | Portuguese | | `ru` | Russian | | `sv` | Swedish | | `nl` | Dutch | | `pl` | Polish | | `it` | Italian | | `da` | Danish | | `uk` | Ukrainian | The following languages are available in beta: | Code | Language | | ----- | ----------------- | | `af` | Afrikaans | | `sq` | Albanian | | `am` | Amharic | | `ar` | Arabic | | `hy` | Armenian | | `as` | Assamese | | `az` | Azerbaijani | | `ba` | Bashkir | | `eu` | Basque | | `be` | Belarusian | | `bn` | Bengali | | `bs` | Bosnian | | `br` | Breton | | `bg` | Bulgarian | | `my` | Burmese | | `yue` | Cantonese | | `ca` | Catalan | | `zh` | Chinese | | `hr` | Croatian | | `cs` | Czech | | `et` | Estonian | | `fo` | Faroese | | `fi` | Finnish | | `gl` | Galician | | `ka` | Georgian | | `el` | Greek | | `gu` | Gujarati | | `ht` | Haitian Creole | | `ha` | Hausa | | `haw` | Hawaiian | | `he` | Hebrew | | `hi` | Hindi | | `hu` | Hungarian | | `is` | Icelandic | | `id` | Indonesian | | `ja` | Japanese | | `jw` | Javanese | | `kn` | Kannada | | `kk` | Kazakh | | `km` | Khmer | | `ko` | Korean | | `lo` | Lao | | `la` | Latin | | `lv` | Latvian | | `ln` | Lingala | | `lt` | Lithuanian | | `lb` | Luxembourgish | | `mk` | Macedonian | | `mg` | Malagasy | | `ms` | Malay | | `ml` | Malayalam | | `mt` | Maltese | | `mi` | Maori | | `mr` | Marathi | | `mn` | Mongolian | | `ne` | Nepali | | `no` | Norwegian | | `nn` | Norwegian Nynorsk | | `oc` | Occitan | | `ps` | Pashto | | `fa` | Persian | | `pa` | Punjabi | | `ro` | Romanian | | `sa` | Sanskrit | | `sr` | Serbian | | `sn` | Shona | | `sd` | Sindhi | | `si` | Sinhala | | `sk` | Slovak | | `sl` | Slovenian | | `so` | Somali | | `su` | Sundanese | | `sw` | Swahili | | `tl` | Tagalog | | `tg` | Tajik | | `ta` | Tamil | | `tt` | Tatar | | `te` | Telugu | | `th` | Thai | | `bo` | Tibetan | | `tr` | Turkish | | `tk` | Turkmen | | `ur` | Urdu | | `uz` | Uzbek | | `vi` | Vietnamese | | `cy` | Welsh | | `yi` | Yiddish | | `yo` | Yoruba | Case does not matter (`DE` is `de`). Region subtags are not accepted: send `pt`, not `pt-BR`. **Send `language` when you know it.** When it is omitted, the model spends an extra decoding pass on the first 30 seconds of audio to detect the language, and a wrong guess (a short clip, a noisy opening, a name in another language) degrades the whole transcript. When you pass it, that pass is skipped and the result is the language you sent. ## The response For a clip in which someone says *"I I want um to go home"*: ```json { "language": "en", "duration": 2.84, "verbatim": { "text": "I I want um to go home", "words": [ { "word": " I", "start": 0.12, "end": 0.26, "alignment_index": 0 }, { "word": " I", "start": 0.34, "end": 0.48, "alignment_index": 1 }, { "word": " want", "start": 0.52, "end": 0.81, "alignment_index": 2 }, { "word": " um", "start": 0.95, "end": 1.31, "alignment_index": 3 }, { "word": " to", "start": 1.58, "end": 1.69, "alignment_index": 4 }, { "word": " go", "start": 1.72, "end": 1.94, "alignment_index": 5 }, { "word": " home", "start": 1.98, "end": 2.41, "alignment_index": 6 } ] }, "intended": { "text": "I want to go home", "words": [ { "word": " I", "start": 0.34, "end": 0.48, "alignment_index": 1 }, { "word": " want", "start": 0.52, "end": 0.81, "alignment_index": 2 }, { "word": " to", "start": 1.58, "end": 1.69, "alignment_index": 4 }, { "word": " go", "start": 1.72, "end": 1.94, "alignment_index": 5 }, { "word": " home", "start": 1.98, "end": 2.41, "alignment_index": 6 } ] }, "alignment": [ { "type": "repetition", "verbatim": "I", "intended": null }, { "type": "match", "verbatim": "I", "intended": "I" }, { "type": "match", "verbatim": "want", "intended": "want" }, { "type": "filler", "verbatim": "um", "intended": null }, { "type": "match", "verbatim": "to", "intended": "to" }, { "type": "match", "verbatim": "go", "intended": "go" }, { "type": "match", "verbatim": "home", "intended": "home" } ], "usage": { "type": "duration", "seconds": 3 } } ``` ### `verbatim` and `intended` Both are a `text` string plus a `words` list, and both come from **one** decode: CrisperWhisper 2 produces the two transcripts together rather than transcribing once and cleaning up afterwards, so they agree on what was in the audio and differ only in what they keep. * `verbatim` is what was said. It keeps fillers (*um*, *uh*), repetitions, false starts, and cut-off words. Use it when the manner of speaking is the point: clinical assessment, speech therapy, fluency research, subtitles that must be faithful. * `intended` is what was meant: the same speech read as fluent text. Use it for notes, search, summaries, or anything downstream that would trip over a stutter. ### Words and timestamps Each `word` has a `start` and `end` in **seconds from the start of the recording**, and keeps its **leading space**. Concatenating the words reproduces the track's `text`: ```python assert "".join(w["word"] for w in track["words"]).strip() == track["text"] ``` The leading space is how a word tells you whether it began a new token or continued one; keep it if you rebuild text, strip it if you display words on their own. Timestamps are per track: each transcript is timed from its own decode, so a word that appears in both tracks (the second `I` above) normally carries the same timing in both, but the two are not guaranteed to be identical to the millisecond. ### `alignment` and `alignment_index` `alignment` is the link between the two tracks, in reading order. Each entry names the token on the verbatim side, the token on the intended side (or `null` if there is none), and how they relate: | `type` | Meaning | | ------------ | ---------------------------------------------------------------------------------------------------- | | `match` | The same word on both sides. | | `repetition` | A repeated word or phrase; only the final instance is kept in `intended`. | | `disfluent` | A false start or restart that does not appear in `intended`. | | `cutoff` | A word the speaker broke off mid-way. | | `filler` | *um*, *uh*, *er* and the like. | | `sound` | A non-word vocalisation: a laugh, a cough, a click. | | `correction` | The speaker corrected themselves; `verbatim` is what was first said, `intended` is what replaced it. | Each word's `alignment_index` is the index of its entry in `alignment`, so you can walk from a verbatim word to its intended counterpart (or learn that it has none) without string matching. It is `null` when the classifier could not prove the link. A `null` is never a guess: the API does not invent links it cannot show. `alignment` itself is `null` when the classifier could not run on this clip, which happens on very long or very dense recordings that exceed its bounds. That is not an error: `verbatim` and `intended` are complete and correct without it, and the request is billed normally. ### Audio without speech Silence, music or noise is a valid request, not an error: both transcripts come back with an empty `text` and an empty `words` list, `alignment` is `[]`, and the audio is billed like any other. Check `verbatim.text` before assuming something was said. ### `usage` `usage.seconds` is the audio's decoded duration rounded **up** to a whole second: what you were charged for, at $0.90 per audio hour ($0.00025 per second). The 2.84 s clip above is 3 seconds, \$0.00075. See [Billing](/billing). ## Python ```python import os import requests with open("clip.wav", "rb") as audio: response = requests.post( "https://api.nyra-labs.com/v1/audio/transcriptions", headers={"Authorization": f"Bearer {os.environ['NYRA_API_KEY']}"}, files={"file": ("clip.wav", audio, "audio/wav")}, data={"model": "crisperwhisper-v2", "language": "en"}, timeout=600, ) response.raise_for_status() result = response.json() print(result["verbatim"]["text"]) print(result["intended"]["text"]) for word in result["verbatim"]["words"]: print(f"{word['start']:6.2f} {word['end']:6.2f} {word['word']}") ``` ## Node ```ts import { openAsBlob } from "node:fs"; const form = new FormData(); form.append("file", await openAsBlob("clip.wav"), "clip.wav"); form.append("model", "crisperwhisper-v2"); form.append("language", "en"); const response = await fetch("https://api.nyra-labs.com/v1/audio/transcriptions", { method: "POST", headers: { Authorization: `Bearer ${process.env.NYRA_API_KEY}` }, body: form, }); if (!response.ok) { throw new Error(`${response.status}: ${await response.text()}`); } const result = await response.json(); console.log(result.verbatim.text); console.log(result.intended.text); ``` ## Limits | | Limit | Over it | | ------------ | ---------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------ | | Upload size | **25 MB** | `413 file_too_large` | | Audio length | **1 hour** of decoded audio | `400 audio_too_long` | | Formats | wav, mp3, m4a, flac, ogg, webm ([details](#accepted-formats)) | `400 unsupported_file_type` (readable audio in another format, or a video); `400 invalid_audio` (not audio at all) | | Requests | 600 per minute per key by default; adjustable per key in the dashboard | `429 rate_limit_exceeded` | Both the size and the length are checked before any money is reserved and before the audio reaches a model. Split long recordings client-side; a 25 MB budget fits about 13 minutes of 16 kHz mono WAV or well over an hour of compressed audio, so which limit you meet first depends on the format. ## Response headers | Header | Meaning | | -------------------------- | ------------------------------------------------------------------------------------------------ | | `x-request-id` | The `req_...` id. Quote it in support requests; the dashboard's request explorer is keyed by it. | | `x-nyra-audio-duration-ms` | The decoded duration, in integer milliseconds. | | `x-nyra-cost-usd` | What this request cost, in USD as a fixed-point decimal. | Rate-limit headers (`x-ratelimit-limit-requests`, `x-ratelimit-remaining-requests`, `x-ratelimit-reset-requests`) are on every response too; see [Authentication](/authentication). ## Retries and idempotency The endpoint accepts an `Idempotency-Key` header of up to 255 characters. Once a request with a given key has completed, a retry with the same key within **24 hours** replays the stored response: the same body and the same `x-request-id`, plus `x-nyra-idempotent-replay: true`, and nothing is charged again. ```bash curl https://api.nyra-labs.com/v1/audio/transcriptions \ -H "Authorization: Bearer $NYRA_API_KEY" \ -H "Idempotency-Key: 6f1c0b4e-2f19-4f0e-b2f1-9a3c5d7e8f01" \ -F "file=@clip.wav" -F "model=crisperwhisper-v2" -F "language=en" ``` Two things it is not: * It is not a lock. Two genuinely concurrent requests with the same key both run and both charge. Wait for the first to finish before retrying. * It is not a cache under zero retention. If your organization stores no transcripts, the response is not kept either; a retry answers `400 idempotency_response_expired` without running or charging again. A key is bound to the model and language of its first request: a retry under the same key with a different `language` or `model` answers `400 idempotency_key_in_use` rather than the other request's transcript. Keys are namespaced to your organization, so nobody else's key collides with yours. ## Errors Every non-2xx response carries the same envelope: ```json { "error": { "message": "The audio is longer than the 3600 second limit. Split it into shorter files.", "type": "invalid_request_error", "param": "file", "code": "audio_too_long" } } ``` `code` is the field to branch on; `message` is for humans and may change. `param` names the offending field when there is one. | Status | `code` | When | | ------ | ------------------------------------ | ------------------------------------------------------------------------------------------------------- | | `400` | `invalid_content_type` | The body is not `multipart/form-data`. | | `400` | `missing_required_parameter` | `file` or `model` was not sent; `param` names it. | | `400` | `invalid_value` | `file` is a text field rather than an upload, or more than one `file` was sent. | | `400` | `invalid_audio` | The file is empty, or its bytes are not audio in any container we can read. | | `400` | `unsupported_file_type` | Readable audio outside the [accepted formats](#accepted-formats), or a file with a video stream. | | `400` | `audio_too_long` | The decoded audio is longer than one hour. | | `400` | `unsupported_language` | `language` is not in the [supported list](#supported-languages). | | `400` | `unsupported_parameter` | A removed field such as `response_format` was sent; `param` names it. | | `400` | `model_not_found` | `model` names a model that does not exist or is retired. | | `400` | `idempotency_response_expired` | See [Retries and idempotency](#retries-and-idempotency). | | `400` | `idempotency_key_in_use` | The same `Idempotency-Key` was used with a different `language` or `model`. | | `401` | `missing_api_key`, `invalid_api_key` | No key, or a key that is unknown, revoked or expired. | | `403` | `insufficient_scope` | The key was narrowed and lacks `transcriptions:write`. | | `413` | `file_too_large` | The upload is over 25 MB. | | `429` | `insufficient_quota` | The wallet cannot cover the request's hold. Retrying will not help; top up. | | `429` | `rate_limit_exceeded` | Too many requests in the current minute. `retry-after` says how long to wait. | | `503` | `capacity_exceeded` | The transcription backend is saturated. Nothing was charged. Retry after the `retry-after` header says. | `insufficient_quota` and `rate_limit_exceeded` share a status code but not a remedy: back off on the second, stop and top up on the first. Check `code`, not the status. The [API reference](/api-reference) is generated from the same OpenAPI spec the server is built from. If a behaviour is not in the spec, it is not part of the public API. > One request, two transcripts: what was said and what was meant, with word timestamps.