A transcript of a conversation is half a transcript if you can't tell who said what. ScribeGrab groups the voices it hears and puts Speaker 1, Speaker 2 in front of every line — automatically, in the text and the subtitles.
No account. 10 recordings a day. Others give you three files a day and call it unlimited. Files up to 60 minutes here, and it resets at midnight.
Plain transcript and subtitles right here. Bleeping, translation, burned-in subtitles and files over 90 MB live in the full tool.
Transcribe audio now →It's also called diarization: not recognising who someone is, but noticing that this voice and that voice are two different people, and marking where each one is talking. The result is a transcript you can read like a script instead of a wall of text.
Every stretch of speech is turned into a voice fingerprint — an ECAPA-TDNN embedding — and those fingerprints are clustered: voices that sit close together are treated as one person. Whichever cluster a line falls into decides its label, and the label is written into the TXT, the SRT and the VTT.
It runs by itself. There's no setting to switch on and nothing to tell it in advance, including how many people are in the room.
It can't know anyone's name. It has never heard your voice before, so it says Speaker 1, not "Sarah": read the first line each person says and replace the labels once in your text editor.
The transcript itself is unaffected either way: if the voices can't be told apart reliably, you simply get the normal transcript without labels.
Speaker labels come out of the same upload as everything else, so a recorded call can come back labelled, with subtitles, and with the swearing bleeped if you need to share it. See also transcribing an interview and a Zoom or Teams recording.
Only need the pages covering one speaker's part, not the whole transcript-turned-PDF? ScanReviver's split PDF pulls out just the page range you need.
By Sam Ridder — I build and run ScribeGrab on my own hardware, on my own. Who I am.
No. If two or more distinct voices are heard, the labels are added by themselves, in the transcript and in both subtitle formats.
No, and no tool can without being told who's who beforehand. You get Speaker 1, Speaker 2 and so on; find the first line of each and replace the labels in your text editor.
Two to eight. Above that the voice grouping isn't reliable enough to be worth showing, so it leaves the labels off rather than guessing wrong.
Either it heard one voice, or it heard more than eight groups, which means the recording was too noisy or too overlapping to cluster confidently. The transcript itself is still complete.
Yes, like everything else here: no account, 10 recordings a day, no watermark, and your file is deleted right after processing.
Yes — automatic speaker identification runs on every transcript with more than one voice, no separate speaker identification software or add-on needed. It's on by default and included in the free tier.
Fully automatic. As soon as more than one voice is detected, the transcription is split into speaker turns with a label for each, so you get transcription with speaker labels out of the box rather than a plugin to install or a setting to find.
Both in one: it's a full transcription tool with diarization (speaker separation) included, not a separate step or a paid add-on.
It happens most on crosstalk or two very similar-sounding voices. The labels are estimated by AI, so give overlapping stretches a quick check and swap the label directly in your text editor — the timestamps make it easy to find the exact spot.
Plain "Speaker 1:", "Speaker 2:" and so on at the start of each turn, in both the transcript and the SRT/VTT subtitles — no special software needed to read them, just a text editor.
Yes, the whole process — upload, diarization, transcript — runs in the browser on our GPUs. Nothing to install, and the file is deleted right after processing.
The model listens for changes in voice as it transcribes and splits the text into turns the moment it detects a different speaker, tagging each turn Speaker 1, Speaker 2 and so on — no manual marking, no separate step.
The model listens for changes in voice and splits the text into turns automatically, tagging each one Speaker 1, Speaker 2 and so on directly in the transcript — no manual marking needed.
It's automatic and usually accurate on clear audio with distinct voices; on crosstalk or very similar-sounding speakers, do a quick check and relabel any turn that's off — no re-upload needed for that.