Quickstart
Sixty seconds from sign-up to your first API call — three commands, nothing to install beyond curl.
Get a key
≈ 30 sCreate an account (e-mail + password, no payment details for the free allowance), then press Create key in the dashboard. The secret is shown once — copy it into your shell:
export KEY=sk_test_… # the key you just created
Billing: a free allowance every month, then prepaid credit — no subscription.
The first call
≈ 20 s
One request, one MP3. Billing is per character, the count
comes back in the X-Characters-Billed header:
curl -X POST https://klangzeile.com/v1/tts \
-H "Authorization: Bearer $KEY" \
-H "Content-Type: application/json" \
-d '{"text": "Every line has a voice.", "voice": "en-heart", "format": "mp3"}' \
-o hello.mp3
Voices: GET https://klangzeile.com/v1/voices. Up to 5,000
characters per request, speed 0.5–2.0, formats mp3, wav,
ogg.
Audio in, protocol out
Transcription is asynchronous: upload, get a job id, fetch
the result when it is done. summary=true adds a structured
summary (title, abstract, outline, decisions, open points) written by
Mistral AI in the EU with zero data retention; summary=false
skips the summary step. When no summary could be produced,
summary is null and summary_error
is "summary_unavailable" — the transcript is still complete
and billed normally. When no speech was found the text is empty and
no_speech is true; minutes are still billed.
Audio shorter than 1 second is refused with 422
AUDIO_TOO_SHORT.
curl -X POST https://klangzeile.com/v1/transcriptions \
-H "Authorization: Bearer $KEY" \
-F "file=@meeting.mp3" -F "language=de" -F "summary=true" \
-F "speakers=true" -F "vocabulary=Maria Spilka"
curl "https://klangzeile.com/v1/transcriptions/tr_…?speaker_names=S1:Maria" -H "Authorization: Bearer $KEY"
mp3, wav, m4a, ogg, webm, flac, aiff, amr, wma, caf; up to 200 MB and 3 h per
file. speakers=true labels who spoke (S1, S2, …;
dialogue, speakers with first words and share);
GET /v1/transcriptions/{id}/speakers previews the labels
without consuming the result, and speaker_names=S1:Anna,S2:Ben
on the final GET replaces them — we never guess names.
Speaker labels are acoustic clusters: a report with many short voices,
or two very similar voices, can yield more or fewer labels than people
— check speakers[].first_words and share before
naming them.
vocabulary (comma-separated names and terms, up to 100;
≤ 2,000 characters in total; multi-word terms are split into single words
and at most 100 words are sent) improves spelling. Speech-to-text, speaker labels and the summary are
produced by Mistral AI in the EU under zero data retention. Billed per
started minute when the job finishes
(minutes_billed in the result). The result is returned
once and deleted; unread results expire after 24 h. Optional
webhook_url (https only): we POST {"id", "status"}
there when the job ends. The webhook is a notification only — it never
carries the transcript, and you must treat it as
unauthenticated: always fetch the result yourself with
GET /v1/transcriptions/{id} and your API key. Signature
verification is a documented follow-up. An
Idempotency-Key does not apply to file uploads (multipart)
or bodies above 1 MiB: a retried upload creates a new job, so keep the
id you get back. timestamps=word gives
word-level timestamps; GET …?format=srt or
vtt returns a subtitle file instead of JSON (42 characters
per line, ≤ 7 s per cue, speaker prefixes when speakers were requested,
names via speaker_names) — it is still one delivery, so
choose JSON or a subtitle format for the single fetch; with
format=json you get words and can build any
format yourself. translate=en or translate=de
translates the transcript segment by segment (timestamps and speakers
kept) and writes the summary in that language;
GET …?format=srt&lang=en (or vtt) delivers
subtitles in the translated language. The translation is produced by
Mistral AI in the EU under zero data retention, like the summary.
translation_error is "not_requested",
"same_language" or "translation_unavailable" —
the transcript is still complete and billed.
No code needed. Turn on "Protocol by e-mail" in your dashboard, then mail
a recording as an attachment to protokoll@klangzeile.com. A few minutes later
you get one mail back with a one-time link to the protocol — summary, who said what,
downloads. Put "en" or "de" in the subject to translate it. Up to 25 MB per mail (about
18 MB of audio, roughly 40 minutes of a voice memo).
Glossary. Save names and terms once — in the dashboard or with
PUT https://klangzeile.com/v1/glossary and {"terms": ["Maria Spilka", "Kubernetes"]}
(GET reads it back, an empty list deletes it). Every transcription, by API or
by e-mail, gets them as a spelling aid; vocabulary sent with a job comes first,
at most 100 terms in total. Same limits as vocabulary: up to 100 terms,
50 characters each, 2,000 characters in total. See it on the landing page:
/#glossary.
When something goes wrong
≈ 10 s| 401 | Key missing, wrong or revoked. |
| 402 | Free allowance used and not enough credit — top up in the dashboard. |
| 422 | Invalid input, e.g. unknown voice or speed out of range. |
| 429 | Rate limit — wait for Retry-After seconds. |
All error codes
(field error_code)
| 401 | UNAUTHORIZED | API key missing, malformed or revoked. |
| 402 | CREDIT_EXHAUSTED | Free allowance used up and not enough credit; top up in the dashboard (top_up_url). |
| 404 | NOT_FOUND | Unknown transcription id. |
| 409 | NOT_READY | Transcription not finished yet (speakers preview). |
| 409 | CANCELLING | DELETE while processing: cancelled after the current step, not billed. |
| 409 | FAILED | Speakers preview of a failed job without a more specific code. |
| 409 | CANCELLED | Speakers preview of a cancelled job. |
| 409 | DECODE_FAILED | Speakers preview: the job failed because the file could not be read as audio. |
| 409 | TRANSCRIPTION_UNAVAILABLE | Speakers preview: the job failed because the speech service was unreachable; upload again later. |
| 409 | AUDIO_LOST | Speakers preview: the recording was lost in a restart; upload again. |
| 410 | RESULT_DELIVERED | Result already delivered and deleted. |
| 410 | RESULT_EXPIRED | Result expired after 24 h and was deleted. |
| 410 | RESULT_DELETED | Deleted on your request. |
| 411 | LENGTH_REQUIRED | Upload without a valid Content-Length. |
| 413 | PAYLOAD_TOO_LARGE | Audio larger than 200 MB. |
| 413 | TEXT_TOO_LONG | TTS text over the character limit. |
| 415 | UNSUPPORTED_MEDIA | Not an audio container we accept. |
| 422 | VALIDATION | Invalid field (e.g. language, timestamps, translate, vocabulary, speaker_names, glossary); detail names it. |
| 422 | AUDIO_TOO_SHORT | Audio shorter than 1 second. |
| 422 | AUDIO_TOO_LONG | Audio longer than 3 hours. |
| 422 | NO_SPEAKERS | speaker_names on a result without speaker labels (upload with speakers=true). |
| 422 | UNKNOWN_SPEAKER | speaker_names names a label the result does not have. |
| 422 | SPEAKER_NAME_COLLISION | speaker_names would give two speakers the same name. |
| 422 | NO_TRANSLATION | lang= asks for a translation the result does not have. |
| 422 | EMPTY_TEXT | TTS: field 'text' is empty. |
| 422 | UNKNOWN_VOICE | TTS: unknown voice; see GET /v1/voices. |
| 422 | UNSUPPORTED_FORMAT | TTS: unsupported audio format. |
| 422 | IDEMPOTENCY_CONFLICT | Idempotency-Key was already used with a different payload. |
| 409 | IDEMPOTENCY_IN_PROGRESS | A request with the same Idempotency-Key is still running; retry after Retry-After seconds. |
| 429 | RATE_LIMITED | Rate limit; wait Retry-After seconds. |
| 500 | INTERNAL | Internal error; quote the request_id. |
| 502 | UPSTREAM_UNAVAILABLE | The processing module is unavailable or failed; retry later. |
| 503 | SERVICE_BUSY | Busy; retry after Retry-After seconds. |
| 504 | UPSTREAM_TIMEOUT | The processing module timed out. |
Errors are RFC 7807 application/problem+json.
Every response carries an X-Request-Id — quote it when you
write to us. Retrying a POST safely: send an
Idempotency-Key header (kept 24 h). Full reference:
API reference.
Text you send for speech is processed in memory only. Transcripts are kept encrypted until you fetch them (24 h at most), then deleted. For billing we keep timestamp, quantity and a SHA-256.