Your audio never leaves your device
Transcription with speaker labels, so you can see who said what
Separate speakers automatically and get a transcript that shows each person's turns. Rename the speakers to real names, and export the result. Free, in your browser, with nothing uploaded.
Free, no signup. First run downloads a model once (about 90 MB for the default), then it is cached. MP3, WAV, M4A, OGG, FLAC. Works best up to about 40 minutes per file.
Who said what, without a cloud service
Speaker labels turn a wall of text into a readable conversation. MeetingScribe adds them by running speaker separation in your browser, so you get a transcript where each turn is attributed to a speaker. Many transcription services charge for this or require an upload. Here it is free and stays on your device.
From labels to real names
- Turn on speaker separation and transcribe as usual.
- The transcript shows generic labels like Speaker 00 and Speaker 01 for each turn.
- Type real names next to those labels and apply them, and the whole transcript updates.
- Export the labeled transcript as text, or as SRT and VTT with the names on each caption.
Getting the best results
Speaker separation is easiest when each person has a clear, distinct voice and there is little overlap. Recordings with one microphone per person, or with speakers taking clear turns, come out best. Very similar voices, lots of crosstalk, or background noise will reduce accuracy, so review and fix the labels in the editor when needed. If separation is not useful for a particular file, you can turn it off for a faster plain transcript.
Free and private
There is no account and no charge. The first run downloads the models once and then caches them. As with everything here, your audio never leaves your device.
Frequently asked questions
- How does speaker separation work here?
- The tool runs a speaker-segmentation model in your browser alongside the transcription, then groups the audio into distinct speakers and labels each turn. You can rename the labels to real names.
- How many speakers can it handle?
- It works best for a small number of voices, roughly one to three. More speakers, heavy crosstalk, or noisy audio make separation harder, which is true of any automatic system.
- Is it perfect?
- No automatic speaker labeling is perfect. Similar voices or overlapping speech can be mislabeled. You can correct labels and text in the in-page editor before exporting.
- Does the audio get uploaded?
- No. Both the transcription and the speaker separation run on your device. Nothing is uploaded.