Speech recognition gets everyday words right and stumbles on the words that matter most to you: company names, product names, people and jargon. Giving the model a short list of those terms, called custom vocabulary or context biasing, fixes many of these errors. It is not magic, and it is worth measuring on your own audio before and after.
Why proper nouns go wrong
A speech recognition model does two things at once. It maps sounds to candidate words, and it uses what it has learned about language to decide which candidate is likely. For common words both signals agree. For a rare name they pull apart: the acoustic signal says something like "Mädchenflohmarkt", but the language model has seen that word rarely or never, so it prefers something more familiar that sounds similar.
The typical errors:
- Phonetic respelling: an unknown name is written as it sounds (a surname like "Berger" becomes "Bärger").
- Substitution by a common word: a product name becomes an ordinary word with similar sounds.
- Wrong split or join: a compound brand name comes back as two words, or two words are glued together.
- Inconsistency: the same name is spelled three ways in one transcript, which breaks search.
- Acronyms and codes: letter sequences and invoice numbers such as RE-2026-0042 come out as words or wrong digits.
German compounds and English loanwords in German speech (and the reverse) make this worse, which matters for mixed-language meetings.
What context biasing does
Context biasing tells the model: these words are likely to appear in this recording. During decoding, candidates that match a listed term get a boost, so when the audio is ambiguous between a common word and your term, your term wins more often.
What it does not do:
- It does not add words that were not spoken. A good implementation only tips close calls.
- It does not fix audio problems. A name mumbled across a bad phone line stays hard to recognise.
- It does not guarantee a spelling. It raises the probability; it does not force it.
- It can over-apply. A term that sounds like a common word may start replacing that word where it does not belong. Short, generic terms are the usual culprits.
How strongly a list is weighted and how it is matched (whole phrases, single words, sub-word pieces) differs between vendors and is often not documented. If you are unsure how your provider treats multi-word terms, test it.
Per-account glossary vs per-request vocabulary
There are two places to put a list, and they serve different purposes.
A per-account glossary holds terms that are always relevant: your company name, product names, the names of the people who are in most meetings. You maintain it once and every job uses it.
Per-request vocabulary holds terms that only matter for one recording: the guest in this interview, the customer in this call, the project codename discussed today.
klangzeile offers both. The glossary is set in the dashboard or with the API:
curl -X PUT https://klangzeile.com/v1/glossary \
-H "Authorization: Bearer $KEY" -H "Content-Type: application/json" \
-d '{"terms": ["Muster GmbH", "Kubernetes", "Anna Berger"]}'
GET /v1/glossary reads it back; an empty list deletes it. It applies to every transcription on the account, by API and by e-mail. Per-request terms go into the vocabulary field of the upload as a comma-separated list, and they come first when both are merged.
Limits, and why they exist
Biasing lists are kept short for a reason: the longer the list, the weaker each entry and the higher the chance of false replacements. On klangzeile the limits are the same for glossary and vocabulary:
- at most 100 terms, 50 characters each, 2,000 characters in total;
- per-request terms first, then the glossary, at most 100 terms combined;
- multi-word terms are split into single words for the speech model, and at most 100 words are sent.
The last point has a practical consequence. "Muster GmbH" is sent as "Muster" and "GmbH". A glossary of 60 two-word names is 120 words, and only the first 100 reach the model. If you have many multi-word names, list the distinctive word ("Berger") rather than the full name, and put per-recording terms in vocabulary so they are always among the first.
Practical rules for a good list:
- Include words the model gets wrong, not words it already gets right.
- Spell each term exactly as you want it in the transcript, including capitalisation.
- Avoid short or generic words that could replace common speech.
- Review the glossary when people or products change.
How to measure the effect
Do not rely on a feeling that "it looks better". Take one representative recording and compare.
- Pick a recording of 5 to 15 minutes with the names and terms you care about, in your usual audio quality.
- Write a reference for the terms: listen and count how often each term is actually spoken. You do not need a full manual transcript for this.
- Run without the list. On klangzeile the glossary applies automatically, so save it first (
GET /v1/glossary), clear it, and upload withoutvocabulary. - Run with the list. Restore the glossary with
PUTor pass the terms invocabulary. - Count for each term: spoken, recognised correctly, misspelled, missed. Also count false insertions, where a term appears but was not said.
- Compare term accuracy: correct occurrences divided by spoken occurrences, before and after.
Each run is a separate job and is billed as such. If you want a broader metric, word error rate (WER) over a fully corrected transcript works too, but it dilutes the effect: fixing ten names in a 3,000-word transcript barely moves WER, while term accuracy shows it clearly.
A worked example
On the klangzeile landing page there is a real before/after from a German interview. Without a glossary, one sentence came back as "Da stand, Mädchenflomag startet demnächst hier". With "Mädchenflohmarkt" and "Flying Circus TV" in the glossary, the same sentence read "Da stand, Mädchenflohmarkt startet demnächst hier". One sentence is an illustration, not a measurement; run the steps above on your own recordings to get numbers you can rely on.
A fictional team at Muster GmbH would do the same with its weekly meeting: a reference count of seven terms, one run without and one with the glossary, and a small table per term. If a term gets worse (false insertions), remove it from the glossary rather than adding more.
FAQ
Does custom vocabulary work in both German and English? The list is passed to the model regardless of language. How well it works depends on the term and the audio; test in each language you use.
Should I add every employee's name? Only those who appear in recordings and get misspelled. A long list of names nobody says weakens the ones that matter.
Is my glossary stored? On klangzeile, yes: it is stored encrypted until you change it or delete the account, and it is sent with each of your jobs. Per-request vocabulary is kept only until the job is done. See GDPR-compliant transcription.
Related: SRT vs WebVTT · Speaker diarization explained · Meeting transcription with a summary API · GDPR-compliant transcription in the EU