klangzeile
Menu

Quickstart

Sixty seconds from sign-up to your first API call — three commands, nothing to install beyond curl.

§ 1

Get a key

≈ 30 s

Create an account (e-mail + password, no payment details for the free allowance), then press Create key in the dashboard. The secret is shown once — copy it into your shell:

export KEY=sk_test_…   # the key you just created

Billing: a free allowance every month, then prepaid credit — no subscription.

§ 2

The first call

≈ 20 s

One request, one MP3. Billing is per character, the count comes back in the X-Characters-Billed header:

curl -X POST https://klangzeile.com/v1/tts \
  -H "Authorization: Bearer $KEY" \
  -H "Content-Type: application/json" \
  -d '{"text": "Every line has a voice.", "voice": "en-heart", "format": "mp3"}' \
  -o hello.mp3

Voices: GET https://klangzeile.com/v1/voices. Up to 5,000 characters per request, speed 0.5–2.0, formats mp3, wav, ogg.

Audio in, protocol out

Transcription is asynchronous: upload, get a job id, fetch the result when it is done. summary=true adds a structured summary (title, abstract, outline, decisions, open points) written by Mistral AI in the EU with zero data retention; summary=false skips the summary step. When no summary could be produced, summary is null and summary_error is "summary_unavailable" — the transcript is still complete and billed normally. When no speech was found the text is empty and no_speech is true; minutes are still billed. Audio shorter than 1 second is refused with 422 AUDIO_TOO_SHORT.

curl -X POST https://klangzeile.com/v1/transcriptions \
  -H "Authorization: Bearer $KEY" \
  -F "file=@meeting.mp3" -F "language=de" -F "summary=true" \
  -F "speakers=true" -F "vocabulary=Maria Spilka"

curl "https://klangzeile.com/v1/transcriptions/tr_…?speaker_names=S1:Maria" -H "Authorization: Bearer $KEY"

mp3, wav, m4a, ogg, webm, flac, aiff, amr, wma, caf; up to 200 MB and 3 h per file. speakers=true labels who spoke (S1, S2, …; dialogue, speakers with first words and share); GET /v1/transcriptions/{id}/speakers previews the labels without consuming the result, and speaker_names=S1:Anna,S2:Ben on the final GET replaces them — we never guess names. Speaker labels are acoustic clusters: a report with many short voices, or two very similar voices, can yield more or fewer labels than people — check speakers[].first_words and share before naming them. vocabulary (comma-separated names and terms, up to 100; ≤ 2,000 characters in total; multi-word terms are split into single words and at most 100 words are sent) improves spelling. Speech-to-text, speaker labels and the summary are produced by Mistral AI in the EU under zero data retention. Billed per started minute when the job finishes (minutes_billed in the result). The result is returned once and deleted; unread results expire after 24 h. Optional webhook_url (https only): we POST {"id", "status"} there when the job ends. The webhook is a notification only — it never carries the transcript, and you must treat it as unauthenticated: always fetch the result yourself with GET /v1/transcriptions/{id} and your API key. Signature verification is a documented follow-up. An Idempotency-Key does not apply to file uploads (multipart) or bodies above 1 MiB: a retried upload creates a new job, so keep the id you get back. timestamps=word gives word-level timestamps; GET …?format=srt or vtt returns a subtitle file instead of JSON (42 characters per line, ≤ 7 s per cue, speaker prefixes when speakers were requested, names via speaker_names) — it is still one delivery, so choose JSON or a subtitle format for the single fetch; with format=json you get words and can build any format yourself. translate=en or translate=de translates the transcript segment by segment (timestamps and speakers kept) and writes the summary in that language; GET …?format=srt&lang=en (or vtt) delivers subtitles in the translated language. The translation is produced by Mistral AI in the EU under zero data retention, like the summary. translation_error is "not_requested", "same_language" or "translation_unavailable" — the transcript is still complete and billed.

No code needed. Turn on "Protocol by e-mail" in your dashboard, then mail a recording as an attachment to protokoll@klangzeile.com. A few minutes later you get one mail back with a one-time link to the protocol — summary, who said what, downloads. Put "en" or "de" in the subject to translate it. Up to 25 MB per mail (about 18 MB of audio, roughly 40 minutes of a voice memo).

Glossary. Save names and terms once — in the dashboard or with PUT https://klangzeile.com/v1/glossary and {"terms": ["Maria Spilka", "Kubernetes"]} (GET reads it back, an empty list deletes it). Every transcription, by API or by e-mail, gets them as a spelling aid; vocabulary sent with a job comes first, at most 100 terms in total. Same limits as vocabulary: up to 100 terms, 50 characters each, 2,000 characters in total. See it on the landing page: /#glossary.

§ 3

When something goes wrong

≈ 10 s
401Key missing, wrong or revoked.
402Free allowance used and not enough credit — top up in the dashboard.
422Invalid input, e.g. unknown voice or speed out of range.
429Rate limit — wait for Retry-After seconds.

All error codes (field error_code)

401UNAUTHORIZEDAPI key missing, malformed or revoked.
402CREDIT_EXHAUSTEDFree allowance used up and not enough credit; top up in the dashboard (top_up_url).
404NOT_FOUNDUnknown transcription id.
409NOT_READYTranscription not finished yet (speakers preview).
409CANCELLINGDELETE while processing: cancelled after the current step, not billed.
409FAILEDSpeakers preview of a failed job without a more specific code.
409CANCELLEDSpeakers preview of a cancelled job.
409DECODE_FAILEDSpeakers preview: the job failed because the file could not be read as audio.
409TRANSCRIPTION_UNAVAILABLESpeakers preview: the job failed because the speech service was unreachable; upload again later.
409AUDIO_LOSTSpeakers preview: the recording was lost in a restart; upload again.
410RESULT_DELIVEREDResult already delivered and deleted.
410RESULT_EXPIREDResult expired after 24 h and was deleted.
410RESULT_DELETEDDeleted on your request.
411LENGTH_REQUIREDUpload without a valid Content-Length.
413PAYLOAD_TOO_LARGEAudio larger than 200 MB.
413TEXT_TOO_LONGTTS text over the character limit.
415UNSUPPORTED_MEDIANot an audio container we accept.
422VALIDATIONInvalid field (e.g. language, timestamps, translate, vocabulary, speaker_names, glossary); detail names it.
422AUDIO_TOO_SHORTAudio shorter than 1 second.
422AUDIO_TOO_LONGAudio longer than 3 hours.
422NO_SPEAKERSspeaker_names on a result without speaker labels (upload with speakers=true).
422UNKNOWN_SPEAKERspeaker_names names a label the result does not have.
422SPEAKER_NAME_COLLISIONspeaker_names would give two speakers the same name.
422NO_TRANSLATIONlang= asks for a translation the result does not have.
422EMPTY_TEXTTTS: field 'text' is empty.
422UNKNOWN_VOICETTS: unknown voice; see GET /v1/voices.
422UNSUPPORTED_FORMATTTS: unsupported audio format.
422IDEMPOTENCY_CONFLICTIdempotency-Key was already used with a different payload.
409IDEMPOTENCY_IN_PROGRESSA request with the same Idempotency-Key is still running; retry after Retry-After seconds.
429RATE_LIMITEDRate limit; wait Retry-After seconds.
500INTERNALInternal error; quote the request_id.
502UPSTREAM_UNAVAILABLEThe processing module is unavailable or failed; retry later.
503SERVICE_BUSYBusy; retry after Retry-After seconds.
504UPSTREAM_TIMEOUTThe processing module timed out.

Errors are RFC 7807 application/problem+json. Every response carries an X-Request-Id — quote it when you write to us. Retrying a POST safely: send an Idempotency-Key header (kept 24 h). Full reference: API reference.

Text you send for speech is processed in memory only. Transcripts are kept encrypted until you fetch them (24 h at most), then deleted. For billing we keep timestamp, quantity and a SHA-256.