> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nyra-labs.com/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nyra-labs.com/_mcp/server.

# Quickstart

`POST /v1/audio/transcriptions` takes an audio file and returns a **verbatim**
transcript (fillers, repetitions and cut-offs included), an **intended**
transcript (the fluent version), word timestamps for both, and an alignment
that links the two.

#### Get a key

Create one in the [dashboard](https://platform.nyra-labs.com/api-keys). The
plaintext is shown once.

```bash
export NYRA_API_KEY="nl_live_YxQFoF6p2xKvStMzWhYhwPMNSlEnWTwzYL83OJCbyTo"
```

#### Top up the wallet

Billing is prepaid. A key with an empty wallet answers `429
    insufficient_quota`. See [Billing](/billing).

#### Transcribe

Send the file as `multipart/form-data`.

```bash
curl https://api.nyra-labs.com/v1/audio/transcriptions \
  -H "Authorization: Bearer $NYRA_API_KEY" \
  -F "file=@clip.wav" \
  -F "model=crisperwhisper-v2" \
  -F "language=en"
```

## Parameters

The request is `multipart/form-data`; any other body answers
`400 invalid_content_type` before it is read.

| Field      |                       | Notes                                                                                                                                                                    |
| ---------- | --------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `file`     | required              | One audio file, up to **25 MB**, in one of the [accepted formats](#accepted-formats). Exactly one `file` part; a text field, a missing part or a second file is a `400`. |
| `model`    | required              | `crisperwhisper-v2`. Any other value answers `400 model_not_found`; `GET /v1/models` lists what is available.                                                            |
| `language` | optional, recommended | ISO-639-1 code from the [supported list](#supported-languages), e.g. `en`, `de`. Anything else answers `400 unsupported_language`.                                       |

### Accepted formats

* `wav`
* `mp3`
* `m4a` (AAC)
* `flac`
* `ogg` (Vorbis or Opus)
* `webm` (audio only)

Any other format, and any file with a video stream, answers
`400 unsupported_file_type`. The model works on 16 kHz mono audio, so to keep
requests no larger than necessary, resample your audio to 16 kHz before you
send it.

### Supported languages

`language` is an ISO-639-1 code (three letters where no two-letter code
exists). These languages are officially supported:

| Code | Language   |
| ---- | ---------- |
| `de` | German     |
| `en` | English    |
| `es` | Spanish    |
| `fr` | French     |
| `pt` | Portuguese |
| `ru` | Russian    |
| `sv` | Swedish    |
| `nl` | Dutch      |
| `pl` | Polish     |
| `it` | Italian    |
| `da` | Danish     |
| `uk` | Ukrainian  |

The following languages are available in beta:

| Code  | Language          |
| ----- | ----------------- |
| `af`  | Afrikaans         |
| `sq`  | Albanian          |
| `am`  | Amharic           |
| `ar`  | Arabic            |
| `hy`  | Armenian          |
| `as`  | Assamese          |
| `az`  | Azerbaijani       |
| `ba`  | Bashkir           |
| `eu`  | Basque            |
| `be`  | Belarusian        |
| `bn`  | Bengali           |
| `bs`  | Bosnian           |
| `br`  | Breton            |
| `bg`  | Bulgarian         |
| `my`  | Burmese           |
| `yue` | Cantonese         |
| `ca`  | Catalan           |
| `zh`  | Chinese           |
| `hr`  | Croatian          |
| `cs`  | Czech             |
| `et`  | Estonian          |
| `fo`  | Faroese           |
| `fi`  | Finnish           |
| `gl`  | Galician          |
| `ka`  | Georgian          |
| `el`  | Greek             |
| `gu`  | Gujarati          |
| `ht`  | Haitian Creole    |
| `ha`  | Hausa             |
| `haw` | Hawaiian          |
| `he`  | Hebrew            |
| `hi`  | Hindi             |
| `hu`  | Hungarian         |
| `is`  | Icelandic         |
| `id`  | Indonesian        |
| `ja`  | Japanese          |
| `jw`  | Javanese          |
| `kn`  | Kannada           |
| `kk`  | Kazakh            |
| `km`  | Khmer             |
| `ko`  | Korean            |
| `lo`  | Lao               |
| `la`  | Latin             |
| `lv`  | Latvian           |
| `ln`  | Lingala           |
| `lt`  | Lithuanian        |
| `lb`  | Luxembourgish     |
| `mk`  | Macedonian        |
| `mg`  | Malagasy          |
| `ms`  | Malay             |
| `ml`  | Malayalam         |
| `mt`  | Maltese           |
| `mi`  | Maori             |
| `mr`  | Marathi           |
| `mn`  | Mongolian         |
| `ne`  | Nepali            |
| `no`  | Norwegian         |
| `nn`  | Norwegian Nynorsk |
| `oc`  | Occitan           |
| `ps`  | Pashto            |
| `fa`  | Persian           |
| `pa`  | Punjabi           |
| `ro`  | Romanian          |
| `sa`  | Sanskrit          |
| `sr`  | Serbian           |
| `sn`  | Shona             |
| `sd`  | Sindhi            |
| `si`  | Sinhala           |
| `sk`  | Slovak            |
| `sl`  | Slovenian         |
| `so`  | Somali            |
| `su`  | Sundanese         |
| `sw`  | Swahili           |
| `tl`  | Tagalog           |
| `tg`  | Tajik             |
| `ta`  | Tamil             |
| `tt`  | Tatar             |
| `te`  | Telugu            |
| `th`  | Thai              |
| `bo`  | Tibetan           |
| `tr`  | Turkish           |
| `tk`  | Turkmen           |
| `ur`  | Urdu              |
| `uz`  | Uzbek             |
| `vi`  | Vietnamese        |
| `cy`  | Welsh             |
| `yi`  | Yiddish           |
| `yo`  | Yoruba            |

Case does not matter (`DE` is `de`). Region subtags are not accepted: send
`pt`, not `pt-BR`.

**Send `language` when you know it.** When it is omitted, the model spends an
extra decoding pass on the first 30 seconds of audio to detect the language,
and a wrong guess (a short clip, a noisy opening, a name in another language)
degrades the whole transcript. When you pass it, that pass is skipped and the
result is the language you sent.

## The response

For a clip in which someone says *"I I want um to go home"*:

```json
{
  "language": "en",
  "duration": 2.84,
  "verbatim": {
    "text": "I I want um to go home",
    "words": [
      { "word": " I",    "start": 0.12, "end": 0.26, "alignment_index": 0 },
      { "word": " I",    "start": 0.34, "end": 0.48, "alignment_index": 1 },
      { "word": " want", "start": 0.52, "end": 0.81, "alignment_index": 2 },
      { "word": " um",   "start": 0.95, "end": 1.31, "alignment_index": 3 },
      { "word": " to",   "start": 1.58, "end": 1.69, "alignment_index": 4 },
      { "word": " go",   "start": 1.72, "end": 1.94, "alignment_index": 5 },
      { "word": " home", "start": 1.98, "end": 2.41, "alignment_index": 6 }
    ]
  },
  "intended": {
    "text": "I want to go home",
    "words": [
      { "word": " I",    "start": 0.34, "end": 0.48, "alignment_index": 1 },
      { "word": " want", "start": 0.52, "end": 0.81, "alignment_index": 2 },
      { "word": " to",   "start": 1.58, "end": 1.69, "alignment_index": 4 },
      { "word": " go",   "start": 1.72, "end": 1.94, "alignment_index": 5 },
      { "word": " home", "start": 1.98, "end": 2.41, "alignment_index": 6 }
    ]
  },
  "alignment": [
    { "type": "repetition", "verbatim": "I",    "intended": null },
    { "type": "match",      "verbatim": "I",    "intended": "I" },
    { "type": "match",      "verbatim": "want", "intended": "want" },
    { "type": "filler",     "verbatim": "um",   "intended": null },
    { "type": "match",      "verbatim": "to",   "intended": "to" },
    { "type": "match",      "verbatim": "go",   "intended": "go" },
    { "type": "match",      "verbatim": "home", "intended": "home" }
  ],
  "usage": { "type": "duration", "seconds": 3 }
}
```

### `verbatim` and `intended`

Both are a `text` string plus a `words` list, and both come from **one**
decode: CrisperWhisper 2 produces the two transcripts together rather than
transcribing once and cleaning up afterwards, so they agree on what was in the
audio and differ only in what they keep.

* `verbatim` is what was said. It keeps fillers (*um*, *uh*), repetitions,
  false starts, and cut-off words. Use it when the manner of speaking is the
  point: clinical assessment, speech therapy, fluency research, subtitles that
  must be faithful.
* `intended` is what was meant: the same speech read as fluent text. Use it
  for notes, search, summaries, or anything downstream that would trip over a
  stutter.

### Words and timestamps

Each `word` has a `start` and `end` in **seconds from the start of the
recording**, and keeps its **leading space**. Concatenating the words
reproduces the track's `text`:

```python
assert "".join(w["word"] for w in track["words"]).strip() == track["text"]
```

The leading space is how a word tells you whether it began a new token or
continued one; keep it if you rebuild text, strip it if you display words on
their own.

Timestamps are per track: each transcript is timed from its own decode, so a
word that appears in both tracks (the second `I` above) normally carries the
same timing in both, but the two are not guaranteed to be identical to the
millisecond.

### `alignment` and `alignment_index`

`alignment` is the link between the two tracks, in reading order. Each entry
names the token on the verbatim side, the token on the intended side (or
`null` if there is none), and how they relate:

| `type`       | Meaning                                                                                              |
| ------------ | ---------------------------------------------------------------------------------------------------- |
| `match`      | The same word on both sides.                                                                         |
| `repetition` | A repeated word or phrase; only the final instance is kept in `intended`.                            |
| `disfluent`  | A false start or restart that does not appear in `intended`.                                         |
| `cutoff`     | A word the speaker broke off mid-way.                                                                |
| `filler`     | *um*, *uh*, *er* and the like.                                                                       |
| `sound`      | A non-word vocalisation: a laugh, a cough, a click.                                                  |
| `correction` | The speaker corrected themselves; `verbatim` is what was first said, `intended` is what replaced it. |

Each word's `alignment_index` is the index of its entry in `alignment`, so
you can walk from a verbatim word to its intended counterpart (or learn that
it has none) without string matching. It is `null` when the classifier could
not prove the link. A `null` is never a guess: the API does not invent links
it cannot show.

`alignment` itself is `null` when the classifier could not run on this clip,
which happens on very long or very dense recordings that exceed its bounds.
That is not an error: `verbatim` and `intended` are complete and correct
without it, and the request is billed normally.

### Audio without speech

Silence, music or noise is a valid request, not an error: both transcripts
come back with an empty `text` and an empty `words` list, `alignment` is `[]`,
and the audio is billed like any other. Check `verbatim.text` before
assuming something was said.

### `usage`

`usage.seconds` is the audio's decoded duration rounded **up** to a whole
second: what you were charged for, at $0.90 per audio hour ($0.00025 per
second). The 2.84 s clip above is 3 seconds, \$0.00075. See
[Billing](/billing).

## Python

```python
import os
import requests

with open("clip.wav", "rb") as audio:
    response = requests.post(
        "https://api.nyra-labs.com/v1/audio/transcriptions",
        headers={"Authorization": f"Bearer {os.environ['NYRA_API_KEY']}"},
        files={"file": ("clip.wav", audio, "audio/wav")},
        data={"model": "crisperwhisper-v2", "language": "en"},
        timeout=600,
    )

response.raise_for_status()
result = response.json()

print(result["verbatim"]["text"])
print(result["intended"]["text"])
for word in result["verbatim"]["words"]:
    print(f"{word['start']:6.2f} {word['end']:6.2f} {word['word']}")
```

## Node

```ts
import { openAsBlob } from "node:fs";

const form = new FormData();
form.append("file", await openAsBlob("clip.wav"), "clip.wav");
form.append("model", "crisperwhisper-v2");
form.append("language", "en");

const response = await fetch("https://api.nyra-labs.com/v1/audio/transcriptions", {
  method: "POST",
  headers: { Authorization: `Bearer ${process.env.NYRA_API_KEY}` },
  body: form,
});

if (!response.ok) {
  throw new Error(`${response.status}: ${await response.text()}`);
}

const result = await response.json();
console.log(result.verbatim.text);
console.log(result.intended.text);
```

## Limits

|              | Limit                                                                  | Over it                                                                                                            |
| ------------ | ---------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------ |
| Upload size  | **25 MB**                                                              | `413 file_too_large`                                                                                               |
| Audio length | **1 hour** of decoded audio                                            | `400 audio_too_long`                                                                                               |
| Formats      | wav, mp3, m4a, flac, ogg, webm ([details](#accepted-formats))          | `400 unsupported_file_type` (readable audio in another format, or a video); `400 invalid_audio` (not audio at all) |
| Requests     | 600 per minute per key by default; adjustable per key in the dashboard | `429 rate_limit_exceeded`                                                                                          |

Both the size and the length are checked before any money is reserved and
before the audio reaches a model. Split long recordings client-side; a 25 MB
budget fits about 13 minutes of 16 kHz mono WAV or well over an hour of
compressed audio, so which limit you meet first depends on the format.

## Response headers

| Header                     | Meaning                                                                                          |
| -------------------------- | ------------------------------------------------------------------------------------------------ |
| `x-request-id`             | The `req_...` id. Quote it in support requests; the dashboard's request explorer is keyed by it. |
| `x-nyra-audio-duration-ms` | The decoded duration, in integer milliseconds.                                                   |
| `x-nyra-cost-usd`          | What this request cost, in USD as a fixed-point decimal.                                         |

Rate-limit headers (`x-ratelimit-limit-requests`, `x-ratelimit-remaining-requests`,
`x-ratelimit-reset-requests`) are on every response too; see
[Authentication](/authentication).

## Retries and idempotency

The endpoint accepts an `Idempotency-Key` header of up to 255 characters.
Once a request with a given key has completed, a retry with the same key
within **24 hours** replays the stored response: the same body and the same
`x-request-id`, plus `x-nyra-idempotent-replay: true`, and nothing is charged
again.

```bash
curl https://api.nyra-labs.com/v1/audio/transcriptions \
  -H "Authorization: Bearer $NYRA_API_KEY" \
  -H "Idempotency-Key: 6f1c0b4e-2f19-4f0e-b2f1-9a3c5d7e8f01" \
  -F "file=@clip.wav" -F "model=crisperwhisper-v2" -F "language=en"
```

Two things it is not:

* It is not a lock. Two genuinely concurrent requests with the same key both
  run and both charge. Wait for the first to finish before retrying.
* It is not a cache under zero retention. If your organization stores no
  transcripts, the response is not kept either; a retry answers
  `400 idempotency_response_expired` without running or charging again.

A key is bound to the model and language of its first request: a retry under
the same key with a different `language` or `model` answers
`400 idempotency_key_in_use` rather than the other request's transcript.

Keys are namespaced to your organization, so nobody else's key collides with
yours.

## Errors

Every non-2xx response carries the same envelope:

```json
{
  "error": {
    "message": "The audio is longer than the 3600 second limit. Split it into shorter files.",
    "type": "invalid_request_error",
    "param": "file",
    "code": "audio_too_long"
  }
}
```

`code` is the field to branch on; `message` is for humans and may change.
`param` names the offending field when there is one.

| Status | `code`                               | When                                                                                                    |
| ------ | ------------------------------------ | ------------------------------------------------------------------------------------------------------- |
| `400`  | `invalid_content_type`               | The body is not `multipart/form-data`.                                                                  |
| `400`  | `missing_required_parameter`         | `file` or `model` was not sent; `param` names it.                                                       |
| `400`  | `invalid_value`                      | `file` is a text field rather than an upload, or more than one `file` was sent.                         |
| `400`  | `invalid_audio`                      | The file is empty, or its bytes are not audio in any container we can read.                             |
| `400`  | `unsupported_file_type`              | Readable audio outside the [accepted formats](#accepted-formats), or a file with a video stream.        |
| `400`  | `audio_too_long`                     | The decoded audio is longer than one hour.                                                              |
| `400`  | `unsupported_language`               | `language` is not in the [supported list](#supported-languages).                                        |
| `400`  | `unsupported_parameter`              | A removed field such as `response_format` was sent; `param` names it.                                   |
| `400`  | `model_not_found`                    | `model` names a model that does not exist or is retired.                                                |
| `400`  | `idempotency_response_expired`       | See [Retries and idempotency](#retries-and-idempotency).                                                |
| `400`  | `idempotency_key_in_use`             | The same `Idempotency-Key` was used with a different `language` or `model`.                             |
| `401`  | `missing_api_key`, `invalid_api_key` | No key, or a key that is unknown, revoked or expired.                                                   |
| `403`  | `insufficient_scope`                 | The key was narrowed and lacks `transcriptions:write`.                                                  |
| `413`  | `file_too_large`                     | The upload is over 25 MB.                                                                               |
| `429`  | `insufficient_quota`                 | The wallet cannot cover the request's hold. Retrying will not help; top up.                             |
| `429`  | `rate_limit_exceeded`                | Too many requests in the current minute. `retry-after` says how long to wait.                           |
| `503`  | `capacity_exceeded`                  | The transcription backend is saturated. Nothing was charged. Retry after the `retry-after` header says. |

`insufficient_quota` and `rate_limit_exceeded` share a status code but not a
remedy: back off on the second, stop and top up on the first. Check `code`,
not the status.

The [API reference](/api-reference) is generated from the same OpenAPI spec
the server is built from. If a behaviour is not in the spec, it is not part
of the public API.