> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://docs.nyra-labs.com/api-reference/audio/transcribe/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nyra-labs.com/_mcp/server. # Create transcription POST https://api.nyra-labs.com/v1/audio/transcriptions Content-Type: multipart/form-data Transcribes a recording and answers with two transcripts: `verbatim`, what was said, including fillers, repetitions and cut-offs; and `intended`, what was meant, as fluent text. Both carry word-level timestamps, and `alignment` relates the two word by word. Send `language` when you know it. The [quickstart](/quickstart) walks through the request, the accepted formats, the [supported languages](/quickstart#supported-languages) and every field of the response. Reference: https://docs.nyra-labs.com/api-reference/audio/transcribe ## Authentication - `Authorization` header (bearer token, required) — Your API key, e.g. `Authorization: Bearer nl_live_...`. ## Servers - `https://api.nyra-labs.com` (Production, default) - `http://localhost:3001` (Local) ## Request ### Body (multipart/form-data) This endpoint expects a multipart form containing a file. - `file` (file, required) — The recording to transcribe, up to 25 MB. Accepted formats, detected from the bytes rather than the filename: wav, mp3, m4a (AAC in an MP4 container), flac, ogg (Vorbis or Opus) and webm. Video files are refused; send the audio track. - `language` (enum, optional) — The language of the audio as an ISO-639-1 code, e.g. `en` or `de`. Optional but recommended: when it is omitted the language is detected from the first 30 seconds, which costs a little time and can guess wrong on short or noisy recordings. See the [supported languages](/quickstart#supported-languages): German, English, Spanish, French, Portuguese, Russian, Swedish, Dutch, Polish, Italian, Danish and Ukrainian are officially supported; the other codes are available in beta. A code outside the list answers `unsupported_language`. - `model` (enum, required) — The model to transcribe with. Required; `crisperwhisper-v2` is the only model, and any other value answers `model_not_found`. ## Response ### 200 The transcription. - `alignment` (list of AlignmentEntry, required, nullable) — The verbatim transcript aligned against the intended one, in reading order. `null` when the two could not be aligned; the transcripts are still valid. - `duration` (double, required) — Duration of the decoded audio, in seconds. - `intended` (Transcript, required) — What the speaker meant: fluent, without fillers, repetitions or cut-offs. - `language` (string, required) — ISO-639-1 code of the audio's language, as sent or as detected. See the [supported languages](/quickstart#supported-languages): twelve are officially supported, the rest are available in beta. - `usage` (Usage, required) - `verbatim` (TranscriptionVerbatim, required) — What the speaker said, including fillers, repetitions and cut-offs. ## Errors ### 400 Bad Request Error The body is not multipart (`invalid_content_type`); `file` or `model` is missing (`missing_required_parameter`); the model does not exist (`model_not_found`); the language is not supported (`unsupported_language`); a removed parameter was sent (`unsupported_parameter`); the file is empty or not audio (`invalid_audio`), not an accepted format or a video (`unsupported_file_type`), or longer than an hour (`audio_too_long`). - `error` (ErrorResponseError, required) ### 401 Unauthorized Error The API key is missing, malformed or invalid. - `error` (ErrorResponseError, required) ### 403 Forbidden Error The API key is missing the `transcriptions:write` scope. - `error` (ErrorResponseError, required) ### 413 Content Too Large Error The upload exceeds the 25 MB limit. - `error` (ErrorResponseError, required) ### 429 Too Many Requests Error Rate limit exceeded, or the organization's wallet balance cannot cover the request (`insufficient_quota`). - `error` (ErrorResponseError, required) ### 503 Service Unavailable Error The transcription backend is at capacity. Retry after the `retry-after` header says. - `error` (ErrorResponseError, required) ## Types ### AlignmentEntry - `intended` (string, required, nullable) — The intended-transcript token, or `null` when nothing was meant here. - `type` (enum, required) — `match`: said and meant. `repetition`: a repeated word. `filler`: um, uh and their kin. `cutoff`: a word broken off. `sound`: a bracketed non-speech event. `disfluent`: verbatim speech with no intended counterpart. `correction`: the verbatim token differs from what was meant. - Allowed values: `match`, `repetition`, `disfluent`, `cutoff`, `filler`, `sound`, `correction` - `verbatim` (string, required, nullable) — The verbatim-transcript token, or `null` when nothing was said here. ### Transcript What the speaker meant: fluent, without fillers, repetitions or cut-offs. - `text` (string, required) - `words` (list of TranscriptWord, required) — Every word, in order, with timings. ### Usage - `seconds` (double, required) — Billed audio seconds: the audio's duration rounded up to a whole second. - `type` (enum, required) - Allowed values: `duration` ### TranscriptionVerbatim What the speaker said, including fillers, repetitions and cut-offs. - `text` (string, required) - `words` (list of TranscriptWord, required) — Every word, in order, with timings. ### ErrorResponseError - `code` (string, required, nullable) - `message` (string, required) - `param` (string, required, nullable) - `type` (enum, required) - Allowed values: `invalid_request_error`, `authentication_error`, `permission_error`, `rate_limit_error`, `insufficient_quota`, `server_error` ### TranscriptWord - `alignment_index` (integer, required, nullable) — Index into `alignment` of the entry this word belongs to, or `null` when the link could not be established. Never guessed. - `end` (double, required) — End of the word, in seconds. - `start` (double, required) — Start of the word, in seconds. - `word` (string, required) — The word, keeping its leading space, so that joining every `word` of a transcript reproduces its `text`. ## Examples **Request** ```json { "file": "", "model": "crisperwhisper-v2" } ``` **Response** ```json { "alignment": [ { "intended": null, "type": "repetition", "verbatim": "I" }, { "intended": "I", "type": "match", "verbatim": "I" }, { "intended": "want", "type": "match", "verbatim": "want" }, { "intended": null, "type": "filler", "verbatim": "um" }, { "intended": "to", "type": "match", "verbatim": "to" }, { "intended": "go", "type": "match", "verbatim": "go" }, { "intended": "home", "type": "match", "verbatim": "home" } ], "duration": 2.84, "intended": { "text": "I want to go home", "words": [ { "alignment_index": 1, "end": 0.48, "start": 0.34, "word": " I" }, { "alignment_index": 2, "end": 0.81, "start": 0.52, "word": " want" }, { "alignment_index": 4, "end": 1.69, "start": 1.58, "word": " to" }, { "alignment_index": 5, "end": 1.94, "start": 1.72, "word": " go" }, { "alignment_index": 6, "end": 2.41, "start": 1.98, "word": " home" } ] }, "language": "en", "usage": { "seconds": 3, "type": "duration" }, "verbatim": { "text": "I I want um to go home", "words": [ { "alignment_index": 0, "end": 0.26, "start": 0.12, "word": " I" }, { "alignment_index": 1, "end": 0.48, "start": 0.34, "word": " I" }, { "alignment_index": 2, "end": 0.81, "start": 0.52, "word": " want" }, { "alignment_index": 3, "end": 1.31, "start": 0.95, "word": " um" }, { "alignment_index": 4, "end": 1.69, "start": 1.58, "word": " to" }, { "alignment_index": 5, "end": 1.94, "start": 1.72, "word": " go" }, { "alignment_index": 6, "end": 2.41, "start": 1.98, "word": " home" } ] } } ``` **SDK Code** ```python import requests url = "https://api.nyra-labs.com/v1/audio/transcriptions" files = { "file": "open('string', 'rb')" } payload = { "language": , "model": "crisperwhisper-v2" } headers = {"Authorization": "Bearer "} response = requests.post(url, data=payload, files=files, headers=headers) print(response.json()) ``` ```javascript const url = 'https://api.nyra-labs.com/v1/audio/transcriptions'; const form = new FormData(); form.append('file', 'string'); form.append('language', ''); form.append('model', 'crisperwhisper-v2'); const options = {method: 'POST', headers: {Authorization: 'Bearer '}}; options.body = form; try { const response = await fetch(url, options); const data = await response.json(); console.log(data); } catch (error) { console.error(error); } ``` ```go package main import ( "fmt" "strings" "net/http" "io" ) func main() { url := "https://api.nyra-labs.com/v1/audio/transcriptions" payload := strings.NewReader("-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"file\"; filename=\"string\"\r\nContent-Type: application/octet-stream\r\n\r\n\r\n-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"language\"\r\n\r\n\r\n-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"model\"\r\n\r\ncrisperwhisper-v2\r\n-----011000010111000001101001--\r\n") req, _ := http.NewRequest("POST", url, payload) req.Header.Add("Authorization", "Bearer ") res, _ := http.DefaultClient.Do(req) defer res.Body.Close() body, _ := io.ReadAll(res.Body) fmt.Println(res) fmt.Println(string(body)) } ``` ```ruby require 'uri' require 'net/http' url = URI("https://api.nyra-labs.com/v1/audio/transcriptions") http = Net::HTTP.new(url.host, url.port) http.use_ssl = true request = Net::HTTP::Post.new(url) request["Authorization"] = 'Bearer ' request.body = "-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"file\"; filename=\"string\"\r\nContent-Type: application/octet-stream\r\n\r\n\r\n-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"language\"\r\n\r\n\r\n-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"model\"\r\n\r\ncrisperwhisper-v2\r\n-----011000010111000001101001--\r\n" response = http.request(request) puts response.read_body ``` ```java import com.mashape.unirest.http.HttpResponse; import com.mashape.unirest.http.Unirest; HttpResponse response = Unirest.post("https://api.nyra-labs.com/v1/audio/transcriptions") .header("Authorization", "Bearer ") .body("-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"file\"; filename=\"string\"\r\nContent-Type: application/octet-stream\r\n\r\n\r\n-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"language\"\r\n\r\n\r\n-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"model\"\r\n\r\ncrisperwhisper-v2\r\n-----011000010111000001101001--\r\n") .asString(); ``` ```php request('POST', 'https://api.nyra-labs.com/v1/audio/transcriptions', [ 'multipart' => [ [ 'name' => 'file', 'filename' => 'string', 'contents' => null ], [ 'name' => 'model', 'contents' => 'crisperwhisper-v2' ] ] 'headers' => [ 'Authorization' => 'Bearer ', ], ]); echo $response->getBody(); ``` ```csharp using RestSharp; var client = new RestClient("https://api.nyra-labs.com/v1/audio/transcriptions"); var request = new RestRequest(Method.POST); request.AddHeader("Authorization", "Bearer "); request.AddParameter("undefined", "-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"file\"; filename=\"string\"\r\nContent-Type: application/octet-stream\r\n\r\n\r\n-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"language\"\r\n\r\n\r\n-----011000010111000001101001\r\nContent-Disposition: form-data; name=\"model\"\r\n\r\ncrisperwhisper-v2\r\n-----011000010111000001101001--\r\n", ParameterType.RequestBody); IRestResponse response = client.Execute(request); ``` ```swift import Foundation let headers = ["Authorization": "Bearer "] let parameters = [ [ "name": "file", "fileName": "string" ], [ "name": "language", "value": ], [ "name": "model", "value": "crisperwhisper-v2" ] ] let boundary = "---011000010111000001101001" var body = "" var error: NSError? = nil for param in parameters { let paramName = param["name"]! body += "--\(boundary)\r\n" body += "Content-Disposition:form-data; name=\"\(paramName)\"" if let filename = param["fileName"] { let contentType = param["content-type"]! let fileContent = String(contentsOfFile: filename, encoding: String.Encoding.utf8) if (error != nil) { print(error as Any) } body += "; filename=\"\(filename)\"\r\n" body += "Content-Type: \(contentType)\r\n\r\n" body += fileContent } else if let paramValue = param["value"] { body += "\r\n\r\n\(paramValue)" } } let request = NSMutableURLRequest(url: NSURL(string: "https://api.nyra-labs.com/v1/audio/transcriptions")! as URL, cachePolicy: .useProtocolCachePolicy, timeoutInterval: 10.0) request.httpMethod = "POST" request.allHTTPHeaderFields = headers request.httpBody = postData as Data let session = URLSession.shared let dataTask = session.dataTask(with: request as URLRequest, completionHandler: { (data, response, error) -> Void in if (error != nil) { print(error as Any) } else { let httpResponse = response as? HTTPURLResponse print(httpResponse) } }) dataTask.resume() ```