Turning a meeting recording into a protocol takes four steps: upload the file, wait for the job, fetch the transcript with its summary once, and optionally translate. This walkthrough shows each step with curl, the error codes you should handle, and how minutes are counted.
Why transcription APIs are asynchronous
A one-hour recording is tens of megabytes and takes longer to process than an HTTP request should stay open. So transcription APIs almost always work in two phases: you upload and immediately get a job id back, then you ask for the result later, either by polling or by receiving a webhook. Designing for that from the start saves you timeouts and duplicate jobs.
The second thing to plan for is retries. A retried upload is a new job unless the API deduplicates it, and in many APIs, klangzeile included, idempotency keys do not apply to multipart uploads. Keep the job id you get back and never re-upload just because a later step failed.
Step 1: upload
The examples below use klangzeile, where one call covers transcript, speaker labels, a structured summary and translation.
curl -X POST https://klangzeile.com/v1/transcriptions \
-H "Authorization: Bearer $KEY" \
-F "file=@weekly.m4a" -F "language=de" \
-F "speakers=true" -F "summary=true" -F "translate=en" \
-F "vocabulary=Muster GmbH,RE-2026-0042"
# → 202 {"id": "tr_…", "status": "queued", "duration_seconds": 2832, …}
What the fields do:
language:de,enorauto. Setting it explicitly avoids a wrong guess on the first seconds of audio.speakers=true: labels S1, S2, … (see speaker diarization explained).summary=true: title, abstract, outline, decisions and open points. It is on by default; sendsummary=falseif you only want the transcript.translate=enorde: translates the transcript segment by segment and writes the summary in that language.vocabulary: names and terms to spell correctly (see custom vocabulary).
Limits: mp3, wav, m4a, ogg, webm, flac and a few more; up to 200 MB and 3 hours per file. Audio shorter than one second is refused.
Step 2: wait
You have two options. With webhook_url (https only) the API sends {"id", "status"} when the job ends. Treat that call as a notification only: it carries no transcript and is not signed, so always fetch the result yourself with your API key.
Without a webhook, poll. There is one trap: the result is delivered once, and a GET on a finished job is the delivery. If you want to set speaker names on the final fetch, poll the speakers preview instead, which does not consume anything:
while :; do
code=$(curl -s -o sp.json -w '%{http_code}' \
"https://klangzeile.com/v1/transcriptions/$ID/speakers" \
-H "Authorization: Bearer $KEY")
[ "$code" = 409 ] && [ "$(jq -r .error_code sp.json)" = NOT_READY ] \
&& { sleep 20; continue; }
break
done
If the loop ends on a 409 with a different error_code, the job failed or was cancelled, and the code says why. Poll every 15 to 30 seconds. Faster polling does not make the job finish sooner and eats into your rate limit, which is lower until the account has any credit (see pricing). The site quotes about two minutes of processing per hour of audio; summary and translation add to that.
Step 3: fetch transcript and summary
Once sp.json shows the speakers, decide who is who and fetch the result:
curl "https://klangzeile.com/v1/transcriptions/$ID?speaker_names=S1:Anna,S2:Ben,S3:Clara" \
-H "Authorization: Bearer $KEY" -o protocol.json
The response contains text, segments with start, end and speaker, a ready-made dialogue, summary with title, abstract, outline, decisions and open_points, the translation, and minutes_billed. Save it immediately. The first successful GET deletes the result on the server side, and unread results expire after 24 hours.
Three fields tell you about partial results. None of them means the transcript is missing:
summary_error:"summary_unavailable"when no summary could be produced;summaryis thennull.translation_error:"not_requested","same_language"or"translation_unavailable".no_speech: truewhen the recording contained no recognisable speech;textis empty.
Step 4: translation and subtitles
With translate=en, the English segments sit in translation.segments with the same timestamps and speakers as the original, and the summary is written in English. If you need subtitles instead of JSON, request ?format=srt&lang=en or format=vtt on the fetch. Remember it is still one delivery: choose JSON or one subtitle file. The trade-offs between the two formats are in SRT vs WebVTT.
Error handling by error_code
Every error is RFC 7807 application/problem+json with a stable error_code and a request_id. Branch on error_code, not on the human-readable detail.
On upload:
| error_code | Status | What to do |
|---|---|---|
UNAUTHORIZED |
401 | Check the key. Do not retry. |
CREDIT_EXHAUSTED |
402 | Free monthly allowance used up and not enough credit; top up in the dashboard (top_up_url). |
PAYLOAD_TOO_LARGE |
413 | Over 200 MB. Re-encode (mono, lower bitrate) or split. |
UNSUPPORTED_MEDIA |
415 | Not an accepted audio container. |
AUDIO_TOO_SHORT, AUDIO_TOO_LONG |
422 | Under 1 second or over 3 hours. |
VALIDATION |
422 | A field is wrong, for example too many vocabulary terms. |
RATE_LIMITED |
429 | Wait for Retry-After seconds, then retry. |
On a job that ends with status: "failed", the GET body carries an error_code such as DECODE_FAILED (the file could not be read as audio) or TRANSCRIPTION_UNAVAILABLE (the speech service was temporarily unreachable; uploading again later is reasonable). On fetch, a 410 means the result is gone: RESULT_DELIVERED, RESULT_EXPIRED or RESULT_DELETED. A 422 with UNKNOWN_SPEAKER or NO_TRANSLATION means your query did not match the result; the result is not consumed, so fix the query and fetch again.
How minutes are billed
- Billing is per started minute of audio, counted when the job finishes. A recording of 47 minutes 12 seconds is 48 minutes.
- Summary, speaker labels, word timestamps and translation do not add minutes.
- A failed or cancelled job costs nothing. A job that finds no speech, or where the summary could not be produced, is billed normally, because the transcription itself ran.
- At upload, the job's price (its minutes, less whatever is left of the free monthly allowance) plus any still-pending jobs' holds are checked against your account: the job is accepted only if the free allowance and credit cover it, and that amount is reserved (held) until the job finishes. Nothing beyond what you have is ever billed; you get a 402
CREDIT_EXHAUSTEDinstead. See pricing for the current numbers.
A worked example
Muster GmbH records its Monday meeting in German, 47 minutes 12 seconds, as weekly.m4a. The operations lead uploads it with translate=en for the English-speaking board member, polls the speakers preview, maps three labels to names, and fetches protocol.json. The protocol lists two decisions (approve the RE-2026-0042 credit note, move the release to next sprint) and one open point (who contacts the supplier). The board member reads the English summary; the team files the German transcript. Billed: 48 minutes.
FAQ
Can I fetch the result twice?
No. Save the first response. If you need JSON and a subtitle file, fetch JSON and build the subtitles from words or segments.
What if my upload times out? Check whether you received a job id before uploading again. Without one, the job may still have been created; a second upload creates and bills a second job.
Is there a no-code option? Yes. With "Protocol by e-mail" switched on in the dashboard you can send a recording to protokoll@klangzeile.com (up to 25 MB) and get a one-time link back.
Related: SRT vs WebVTT · Speaker diarization explained · Custom vocabulary for speech recognition · GDPR-compliant transcription in the EU