MMeetingScribe.net

Your audio never leaves your device

How to correct wrong speaker labels

Automatic speaker labels go wrong in a few specific ways, and each has a different fix. This page explains what causes each one and what actually helps, including measurements from a real four-person recording rather than estimates.

Drop audio here, or click to choose
MP3, WAV, M4A, OGG, FLAC, AMR. Everything runs on your device. Your audio is never uploaded.
Got a long recording in several files? Select them all at once. Name them in order, like meeting_part1.mp3 and meeting_part2.mp3, and they will be joined into one transcript.

Takes a little extra time, and your audio still never leaves your device. Labels go by how voices sound, so treat them as a best effort: one person can be split across two or three labels, similar voices can be merged, and someone joining partway through may be folded into a speaker already talking. Giving a number usually helps. Check them before relying on them.

Specifying this number improves accuracy. Count people who really take part, whenever they join. Someone who only says a word or two is better left out.

Timestamps each word instead of each phrase, so short interjections like "Yeah" go to the right person. On a busy four-person recording this corrected about 1 word in 5; it changes little when people talk in long turns. Adds roughly 40% to the time on a computer, and almost nothing on a phone.

Phone audio loses the high frequencies that separate consonants, so they are put back before transcribing. Only phone-quality recordings are treated, and you are told when it happens. Also even out volume is for one case: someone far from the microphone and much quieter throughout. It cost accuracy on other recordings, so use it only if a quiet person came out badly.

Worth filling in. The speech model knows ordinary English, not your subject, so a word it has never heard becomes whatever sounds closest, and it usually gets that word wrong every single time it comes up. A sentence is enough, and the names and jargon in it are picked out and watched for. It cannot guess a word you do not mention, so name the people and the terms that matter.

Free, no signup. The first time you use it, it takes a minute or two to set up. After that setup is quick. Works with MP3, WAV, M4A, OGG, FLAC and AMR, and long recordings are handled in parts.

Four things that go wrong, and which fix applies

Speaker labels fail in recognisable ways. Working out which one you are looking at matters, because the fixes are not interchangeable and the wrong one can make things worse.

  • Short interjections on the wrong person.A "Yeah" or "Right" is credited to whoever was talking around it. Fixed by word-level timing.
  • One person split into two or three labels. The same voice appears as SPEAKER_01 early and SPEAKER_03 later. Usually fixed by giving a speaker count, otherwise by renaming both to one name.
  • Two people merged under one label. Similar voices, often same gender and similar recording conditions. Hardest to fix after the fact.
  • A late joiner folded into an existing speaker. Someone who arrives partway through gets absorbed into whoever was already talking. Also hard to fix after the fact.

Fix one: time each word instead of each phrase

This is the fix with the clearest evidence behind it, and it addresses the most visible annoyance, which is short interjections landing on the wrong person.

Speech recognition normally produces a timestamp for a whole phrase. Speaker separation produces its own timeline of who is talking when. When those two timelines disagree, and a phrase spans a change of speaker, the phrase can only be given one label, so it goes to whoever holds the larger share of it. Everything else in that phrase is then wrong.

Switching on Sharper speaker labelstimes every word separately, so each word is matched to the speaker talking at that moment rather than inheriting the phrase's label. On a real 19-minute recording with four participants, we measured:

  • 38% of phrases straddled a speaker change, so they carried at least one word belonging to somebody else.
  • 19.6% of words were given a different speaker once timed individually.
  • 1.4 times the transcription time on a computer, and no measurable cost on a phone.

Two caveats worth stating plainly. That recording was unusually conversational, with a speaker change roughly every two seconds, so it sits at the high end of the benefit; a lecture or a one-to-one interview with long turns will gain much less. And the wording of the transcript can differ slightly between the two settings, because asking for word timings changes how the speech model is constrained while it decodes, so this is not purely a relabelling of identical text.

Fix two: say how many people are talking

Left to itself the tool has to work out how many distinct voices exist, and that decision is where most splitting and merging happens. Telling it the number usually improves matters, particularly for the case where one person has been split across several labels.

The important part is how you count. Count people who genuinely take part in the conversation, whenever they join. Do not count someone who says a word or two in passing. Asking for more speakers than really participate does not surface a quiet person: it forces the tool to split a real speaker in two to reach your number. We confirmed this on a real recording, where requesting one more speaker than the tool had found split an existing speaker rather than revealing the missing one.

If you are unsure, an estimate on the low side is safer than one on the high side.

Fix three: rename, and merge by renaming

Renaming is the last step for every transcript, and it doubles as a repair tool. Enter real names against the detected speakers and apply them, and every download afterwards carries the names rather than the placeholders.

The useful trick: give two labels the same name to merge them. If one person was split into SPEAKER_01 and SPEAKER_03, naming both "Alice" produces a transcript with a single consistent Alice, without editing anything by hand. It is the fastest repair for splitting when a speaker count has not solved it.

Renaming rewrites the transcript, the SRT and the VTT together, so do it before you export rather than editing files afterwards. The exact output of each format is on the speaker label format page.

What cannot be repaired afterwards

Being straight about the limits saves you time. Two people merged under one label, and a late joiner absorbed into an existing speaker, are both cases where the evidence needed to separate them is not in the labels. No amount of renaming recovers it, because the tool never made the distinction in the first place. If the passage matters, the practical options are to re-run with a speaker count and see whether the separation changes, or to correct those passages by hand in the transcript editor before exporting.

Speaker separation is a best-effort feature and worth checking before you rely on it. It is strongest with a small number of clearly different voices and a clean recording, and it struggles with crosstalk, similar voices, and anyone who is quiet or distant from the microphone. The speaker labels page covers how it works and how to record so that it works better next time.

A quick way to check

Rather than reading the whole transcript, spot-check the boundaries, because that is where errors concentrate. Look at the first and last few words of several turns. Words stranded on the wrong side of a change of speaker are the signature of phrase-level timing, and they are the ones word-level timing fixes. If instead you find whole passages under a plausible but wrong name, you are looking at a splitting or merging problem, and the speaker count and renaming are your tools.

Try it on your own recording

All of this runs in your browser. Your audio is never uploaded, there is no account, and the transcript, subtitles and speaker labels stay on your device. Open the free transcriber, turn on speaker labels, and if the recording has people talking over each other, turn on Sharper speaker labels as well.

Last reviewed: August 2026

Frequently asked questions

Why is a short 'Yeah' credited to the wrong person?
Because speech recognition normally timestamps a whole phrase rather than each word. If a phrase spans a change of speaker, the entire phrase is credited to whoever holds most of it, and a short interjection sitting inside it goes to the wrong person. Turning on Sharper speaker labels times each word separately, which fixes this specific class of error.
How much difference does word-level timing actually make?
Measured on a real 19-minute four-person conversation: 38 percent of transcript phrases straddled a speaker change, and 19.6 percent of words came out with a different speaker label once each word was timed separately. That file was unusually conversational, with a speaker change roughly every two seconds, so it is towards the high end. A recording where people speak in long uninterrupted turns has far fewer boundaries and gains much less.
Does turning it on make transcription slower?
On a computer it adds roughly 40 percent to the transcription time. On a phone it costs almost nothing, because the extra work is in moving data back from the graphics card and phones do not take that path. The transcript wording can differ slightly between the two, because asking for word timings changes how the speech model is constrained while it decodes.
One person is showing up as two different speakers. What do I do?
This is the most common failure. Voices vary across a long recording, and the tool can decide one person is two. Setting the number of people usually pulls the split back together. If it does not, rename both labels to the same name before exporting, which merges them in the output.
Should I always set the number of speakers?
It usually helps, but count only people who really take part. Counting someone who says a word or two forces the tool to split a genuine speaker in two in order to reach your number, which is worse than leaving that person out. This was confirmed on a real recording: asking for one more speaker than the tool found split an existing speaker rather than surfacing the missing one.
Can it work out who people are by name?
No, and deliberately so. It compares voices only within the single recording you gave it, and holds no voiceprints, no database, and nothing that persists between files. It can tell two voices apart without any idea who they belong to, which is why you supply the names.

Related