SCRIBEGRABOpen the tool →

Identify speakers in audio free

By ShinobiTools Team · Last updated: July 2026

A transcript of a conversation is half a transcript if you can't tell who said what. ScribeGrab groups the voices it hears and puts Speaker 1, Speaker 2 in front of every line — automatically, in the text and the subtitles.

Transcribe audio now →

What speaker identification actually gives you

It's also called diarization: not recognising who someone is, but noticing that this voice and that voice are two different people, and marking where each one is talking. The result is a transcript you can read like a script instead of a wall of text.

How it works here

Every stretch of speech is turned into a voice fingerprint — an ECAPA-TDNN embedding — and those fingerprints are clustered: voices that sit close together are treated as one person. Whichever cluster a line falls into decides its label, and the label is written into the TXT, the SRT and the VTT.

It runs by itself. There's no setting to switch on and nothing to tell it in advance, including how many people are in the room.

Where it stops — honestly

It can't know anyone's name. It has never heard your voice before, so it says Speaker 1, not "Sarah": read the first line each person says and replace the labels once in your text editor.

The transcript itself is unaffected either way: if the voices can't be told apart reliably, you simply get the normal transcript without labels.

Good with

Speaker labels come out of the same upload as everything else, so a recorded call can come back labelled, with subtitles, and with the swearing bleeped if you need to share it. See also transcribing an interview and a Zoom or Teams recording.

Common questions

Do I have to switch anything on?

No. If two or more distinct voices are heard, the labels are added by themselves, in the transcript and in both subtitle formats.

Will it use people's real names?

No, and no tool can without being told who's who beforehand. You get Speaker 1, Speaker 2 and so on; find the first line of each and replace the labels in your text editor.

How many speakers can it handle?

Two to eight. Above that the voice grouping isn't reliable enough to be worth showing, so it leaves the labels off rather than guessing wrong.

Why does my transcript have no labels?

Either it heard one voice, or it heard more than eight groups, which means the recording was too noisy or too overlapping to cluster confidently. The transcript itself is still complete.

Is speaker identification free?

Yes, like everything else here: no account, no daily cap, no watermark, and your file is deleted right after processing.