By ShinobiTools Team · Last updated: July 2026
A transcript of a conversation is half a transcript if you can't tell who said what. ScribeGrab groups the voices it hears and puts Speaker 1, Speaker 2 in front of every line — automatically, in the text and the subtitles.
Transcribe audio now →It's also called diarization: not recognising who someone is, but noticing that this voice and that voice are two different people, and marking where each one is talking. The result is a transcript you can read like a script instead of a wall of text.
Every stretch of speech is turned into a voice fingerprint — an ECAPA-TDNN embedding — and those fingerprints are clustered: voices that sit close together are treated as one person. Whichever cluster a line falls into decides its label, and the label is written into the TXT, the SRT and the VTT.
It runs by itself. There's no setting to switch on and nothing to tell it in advance, including how many people are in the room.
It can't know anyone's name. It has never heard your voice before, so it says Speaker 1, not "Sarah": read the first line each person says and replace the labels once in your text editor.
The transcript itself is unaffected either way: if the voices can't be told apart reliably, you simply get the normal transcript without labels.
Speaker labels come out of the same upload as everything else, so a recorded call can come back labelled, with subtitles, and with the swearing bleeped if you need to share it. See also transcribing an interview and a Zoom or Teams recording.
No. If two or more distinct voices are heard, the labels are added by themselves, in the transcript and in both subtitle formats.
No, and no tool can without being told who's who beforehand. You get Speaker 1, Speaker 2 and so on; find the first line of each and replace the labels in your text editor.
Two to eight. Above that the voice grouping isn't reliable enough to be worth showing, so it leaves the labels off rather than guessing wrong.
Either it heard one voice, or it heard more than eight groups, which means the recording was too noisy or too overlapping to cluster confidently. The transcript itself is still complete.
Yes, like everything else here: no account, no daily cap, no watermark, and your file is deleted right after processing.