klangzeile
Menu

Blog

From recording to protocol: meeting transcription with a summary API

Turning a meeting recording into a protocol takes four steps: upload the file, wait for the job, fetch the transcript with its summary once, and optionally translate. This walkthrough shows each step with curl, the error codes you should handle, and how minutes are counted.

Why transcription APIs are asynchronous

A one-hour recording is tens of megabytes and takes longer to process than an HTTP request should stay open. So transcription APIs almost always work in two phases: you upload and immediately get a job id back, then you ask for the result later, either by polling or by receiving a webhook. Designing for that from the start saves you timeouts and duplicate jobs.

The second thing to plan for is retries. A retried upload is a new job unless the API deduplicates it, and in many APIs, klangzeile included, idempotency keys do not apply to multipart uploads. Keep the job id you get back and never re-upload just because a later step failed.

Step 1: upload

The examples below use klangzeile, where one call covers transcript, speaker labels, a structured summary and translation.

curl -X POST https://klangzeile.com/v1/transcriptions \
  -H "Authorization: Bearer $KEY" \
  -F "file=@weekly.m4a" -F "language=de" \
  -F "speakers=true" -F "summary=true" -F "translate=en" \
  -F "vocabulary=Muster GmbH,RE-2026-0042"
# → 202 {"id": "tr_…", "status": "queued", "duration_seconds": 2832, …}

What the fields do:

  • language: de, en or auto. Setting it explicitly avoids a wrong guess on the first seconds of audio.
  • speakers=true: labels S1, S2, … (see speaker diarization explained).
  • summary=true: title, abstract, outline, decisions and open points. It is on by default; send summary=false if you only want the transcript.
  • translate=en or de: translates the transcript segment by segment and writes the summary in that language.
  • vocabulary: names and terms to spell correctly (see custom vocabulary).

Limits: mp3, wav, m4a, ogg, webm, flac and a few more; up to 200 MB and 3 hours per file. Audio shorter than one second is refused.

Step 2: wait

You have two options. With webhook_url (https only) the API sends {"id", "status"} when the job ends. Treat that call as a notification only: it carries no transcript and is not signed, so always fetch the result yourself with your API key.

Without a webhook, poll. There is one trap: the result is delivered once, and a GET on a finished job is the delivery. If you want to set speaker names on the final fetch, poll the speakers preview instead, which does not consume anything:

while :; do
  code=$(curl -s -o sp.json -w '%{http_code}' \
    "https://klangzeile.com/v1/transcriptions/$ID/speakers" \
    -H "Authorization: Bearer $KEY")
  [ "$code" = 409 ] && [ "$(jq -r .error_code sp.json)" = NOT_READY ] \
    && { sleep 20; continue; }
  break
done

If the loop ends on a 409 with a different error_code, the job failed or was cancelled, and the code says why. Poll every 15 to 30 seconds. Faster polling does not make the job finish sooner and eats into your rate limit, which is lower until the account has any credit (see pricing). The site quotes about two minutes of processing per hour of audio; summary and translation add to that.

Step 3: fetch transcript and summary

Once sp.json shows the speakers, decide who is who and fetch the result:

curl "https://klangzeile.com/v1/transcriptions/$ID?speaker_names=S1:Anna,S2:Ben,S3:Clara" \
  -H "Authorization: Bearer $KEY" -o protocol.json

The response contains text, segments with start, end and speaker, a ready-made dialogue, summary with title, abstract, outline, decisions and open_points, the translation, and minutes_billed. Save it immediately. The first successful GET deletes the result on the server side, and unread results expire after 24 hours.

Three fields tell you about partial results. None of them means the transcript is missing:

  • summary_error: "summary_unavailable" when no summary could be produced; summary is then null.
  • translation_error: "not_requested", "same_language" or "translation_unavailable".
  • no_speech: true when the recording contained no recognisable speech; text is empty.

Step 4: translation and subtitles

With translate=en, the English segments sit in translation.segments with the same timestamps and speakers as the original, and the summary is written in English. If you need subtitles instead of JSON, request ?format=srt&lang=en or format=vtt on the fetch. Remember it is still one delivery: choose JSON or one subtitle file. The trade-offs between the two formats are in SRT vs WebVTT.

Error handling by error_code

Every error is RFC 7807 application/problem+json with a stable error_code and a request_id. Branch on error_code, not on the human-readable detail.

On upload:

error_code Status What to do
UNAUTHORIZED 401 Check the key. Do not retry.
CREDIT_EXHAUSTED 402 Free monthly allowance used up and not enough credit; top up in the dashboard (top_up_url).
PAYLOAD_TOO_LARGE 413 Over 200 MB. Re-encode (mono, lower bitrate) or split.
UNSUPPORTED_MEDIA 415 Not an accepted audio container.
AUDIO_TOO_SHORT, AUDIO_TOO_LONG 422 Under 1 second or over 3 hours.
VALIDATION 422 A field is wrong, for example too many vocabulary terms.
RATE_LIMITED 429 Wait for Retry-After seconds, then retry.

On a job that ends with status: "failed", the GET body carries an error_code such as DECODE_FAILED (the file could not be read as audio) or TRANSCRIPTION_UNAVAILABLE (the speech service was temporarily unreachable; uploading again later is reasonable). On fetch, a 410 means the result is gone: RESULT_DELIVERED, RESULT_EXPIRED or RESULT_DELETED. A 422 with UNKNOWN_SPEAKER or NO_TRANSLATION means your query did not match the result; the result is not consumed, so fix the query and fetch again.

How minutes are billed

  • Billing is per started minute of audio, counted when the job finishes. A recording of 47 minutes 12 seconds is 48 minutes.
  • Summary, speaker labels, word timestamps and translation do not add minutes.
  • A failed or cancelled job costs nothing. A job that finds no speech, or where the summary could not be produced, is billed normally, because the transcription itself ran.
  • At upload, the job's price (its minutes, less whatever is left of the free monthly allowance) plus any still-pending jobs' holds are checked against your account: the job is accepted only if the free allowance and credit cover it, and that amount is reserved (held) until the job finishes. Nothing beyond what you have is ever billed; you get a 402 CREDIT_EXHAUSTED instead. See pricing for the current numbers.

A worked example

Muster GmbH records its Monday meeting in German, 47 minutes 12 seconds, as weekly.m4a. The operations lead uploads it with translate=en for the English-speaking board member, polls the speakers preview, maps three labels to names, and fetches protocol.json. The protocol lists two decisions (approve the RE-2026-0042 credit note, move the release to next sprint) and one open point (who contacts the supplier). The board member reads the English summary; the team files the German transcript. Billed: 48 minutes.

FAQ

Can I fetch the result twice? No. Save the first response. If you need JSON and a subtitle file, fetch JSON and build the subtitles from words or segments.

What if my upload times out? Check whether you received a job id before uploading again. Without one, the job may still have been created; a second upload creates and bills a second job.

Is there a no-code option? Yes. With "Protocol by e-mail" switched on in the dashboard you can send a recording to protokoll@klangzeile.com (up to 25 MB) and get a one-time link back.

Related: SRT vs WebVTT · Speaker diarization explained · Custom vocabulary for speech recognition · GDPR-compliant transcription in the EU

← All articles