Speaker diarizationspeaker labels
Speaker diarization is the process of working out who spoke when in a recording, so a transcript can be labeled by speaker. It answers "who said this?" rather than "what was said?".
Diarization is a separate model from speech recognition. It listens for changes in voice, clusters the segments that sound like the same person, and assigns each a label, typically Speaker 1 and Speaker 2 until a name is attached. Combined with a transcript it produces the familiar meeting-notes layout, with each line attributed.
It is a transcription feature, not a dictation one. When you dictate, there is one speaker and it is you, so a dictation tool has no use for diarization and should not be judged for lacking it. When you transcribe a meeting, an interview, or a podcast, it is the difference between a transcript and a wall of text.
Why it is hard
Two people with similar voices, one person whose voice changes with a cold, overlapping speech, and a laptop microphone that flattens everyone into the same room tone all defeat the clustering. Speaker counts are usually guessed rather than known. The best systems are good on clean, turn-taking audio like a podcast and much worse on a real meeting where three people talk over each other. Expect to fix labels by hand.
Which tools have it
Meeting transcribers like Otter.ai build their product around it. File transcribers such as MacWhisper offer speaker labels on recordings. Dictation tools generally do not, because the job never has two speakers. If diarization is on your list of requirements, you are shopping for a transcriber.
In Flit
Flit does not diarize, because it never records more than one person: you, while you hold the key. For meetings, use a transcriber.
Flit vs Otter.ai, the meeting transcriber →Questions
Fair questions.
What does diarization mean in transcription?
Labeling a transcript by who was speaking. A diarization model detects voice changes and groups segments by speaker, so a meeting transcript reads as a conversation rather than a single block.
Does dictation software do speaker labeling?
Almost never, because dictation has one speaker. Speaker labels are a feature of meeting and file transcription tools. If you need them, choose a transcriber; if you need live dictation, a dictation app.
Where this shows up on flit.fyi
Compared on it
Related terms