Speaker diarization answers "who spoke when" without knowing who anyone is. It groups stretches of audio by voice and labels the groups S1, S2, S3. Turning those labels into names is a separate step that you, not the model, should do, because only you know who was in the room.
What diarization actually does
A transcript tells you what was said. Diarization adds a speaker label to each part of it. Implementations differ, and vendors rarely publish the details of their models, but the classic pipeline has three stages:
- Find speech. Voice activity detection separates speech from silence, music and noise.
- Describe each stretch of speech. Short windows of audio are turned into a vector (an "embedding") that captures what the voice sounds like: pitch, timbre, speaking style, and also the microphone and the room.
- Cluster. Windows with similar embeddings are grouped. Each group becomes one label. The number of groups is usually estimated from the audio, not given in advance.
Newer end-to-end models fold these stages into one network, and some can handle two people talking at once better than the classic pipeline. The output looks the same either way: time ranges with a label.
The key point is that a label is an acoustic cluster, not an identity. S1 means "the voice that sounds like this", not "the chairperson" and not "Anna".
What it can do well
- Separate two to a handful of speakers with distinct voices in a reasonably quiet recording.
- Follow turn-taking in interviews, podcasts and meetings where people let each other finish.
- Give you each speaker's share of talking time and their first words, which is often enough to recognise them.
Where it goes wrong
These are properties of the problem, not of one product. Expect them from any diarization system and check the output accordingly.
- Overlapping speech. When two people talk at once, most systems assign the overlap to one of them. The other person's words are either lost from the transcript or attributed to the wrong speaker.
- Short interjections. "Yes", "right", "mhm" are too short to produce a reliable voice embedding. They often end up with whoever spoke just before.
- Similar voices. Two colleagues of similar age and accent, recorded on the same laptop microphone, can merge into one label.
- One person, two labels. The opposite also happens: someone who moves away from the microphone, gets emotional, or joins by phone halfway through can be split into two clusters.
- Many short voices. A news report or a vox pop with a dozen people saying one sentence each can produce more or fewer labels than there were people.
- Channel effects. In a hybrid meeting, everyone on the video call comes through the same loudspeaker. The system may hear "the loudspeaker" as one voice.
What helps: one microphone per person where possible, a quiet room, people not talking over each other, and a recording format that is not heavily compressed.
Why APIs return S1 and S2 instead of names
A speech model has no way of knowing that the second voice belongs to Ben unless someone told it. Some tools guess from what is said ("Thanks, Ben") and fill in names. That is convenient when it works and quietly wrong when it does not: a transcript that attributes a decision to the wrong named person is worse than one that says S2.
There is a privacy side too. Recognising a specific person by their voice can count as processing biometric data, which the GDPR treats as a special category (Art. 9). Anonymous cluster labels are not meant to identify anyone, which keeps you further from that line. You assign names from your own knowledge of who attended. Where exactly the legal line runs is a question for your data protection officer, not for this article.
Mapping names afterwards
The workflow is: look at what each label said, decide who it is, then replace the labels everywhere, including the summary.
klangzeile splits this into two calls because results are delivered once. GET /v1/transcriptions/{id}/speakers returns the labels with each speaker's first words and share of speech, without consuming the result. The final GET then takes speaker_names=S1:Anna,S2:Ben and replaces the labels in the segments, the dialogue, the speakers block and the summary text. A typo in a label costs a 422 (UNKNOWN_SPEAKER), not the transcript.
One constraint: two labels cannot be given the same name (SPEAKER_NAME_COLLISION). If one person was split into S1 and S3, fetch the JSON with neutral labels and merge them in your own code.
A worked example
Muster GmbH records a 25-minute supplier call with three participants: Anna (purchasing), Ben (finance) and a supplier representative. They upload with speakers=true:
curl -X POST https://klangzeile.com/v1/transcriptions \
-H "Authorization: Bearer $KEY" \
-F "file=@supplier-call.m4a" -F "language=en" -F "speakers=true"
curl https://klangzeile.com/v1/transcriptions/tr_…/speakers \
-H "Authorization: Bearer $KEY"
The preview (fictional values) shows four labels, not three:
{"speaker_count": 4,
"speakers": {
"S1": {"first_words": "Good morning, thanks for making time", "share": 0.41},
"S2": {"first_words": "Morning. I have the figures for RE-2026-0042", "share": 0.22},
"S3": {"first_words": "Hello, can you hear me now", "share": 0.30},
"S4": {"first_words": "Yes.", "share": 0.07}}}
Reading it:
- S1 opens the call, so it is Anna, who organised it.
- S2 talks about the invoice, so it is Ben.
- S3 is the supplier, who joined by phone.
- S4 has 7 percent and starts with "Yes." It is either the supplier after a change in line quality or short interjections split off from someone else.
The team checks two segments of S4 against the audio. Both are the supplier. Because two labels cannot share a name, they fetch the JSON with speaker_names=S1:Anna,S2:Ben,S3:Supplier and relabel S4 in their own script before filing the protocol. Ten minutes of checking saved a transcript that would otherwise show a fourth participant who never existed.
Checking the result
Before you trust speaker labels for anything that matters (minutes, quotes, legal notes):
- Compare
speaker_countwith the number of people who actually spoke. - Look at labels with a very small share; they are often fragments.
- Listen to the first few seconds of each label.
- Read the decisions in the summary with names in place and ask whether the right person is credited.
FAQ
Can diarization tell me who someone is? No. It only tells you that two stretches of audio sound like the same voice. Identifying people would need enrolled voice samples, which is a different technology with different legal requirements.
Does a higher-quality recording fix overlaps? It helps with similar voices and channel effects. Overlapping speech stays hard even in good recordings.
Do speaker labels change the price? With klangzeile, no; billing is per started minute of audio either way. Other vendors may price it differently, so check.
Related: SRT vs WebVTT · Meeting transcription with a summary API · Custom vocabulary for speech recognition · GDPR-compliant transcription in the EU