SCRIBEGRABOpen the tool →

Speaker diarization free, speaker labels in your audio

A transcript of a conversation is half a transcript if you can't tell who said what. ScribeGrab groups the voices it hears and puts Speaker 1, Speaker 2 in front of every line — automatically, in the text and the subtitles.

No account. 10 recordings a day. Others give you three files a day and call it unlimited. Files up to 60 minutes here, and it resets at midnight.

Plain transcript and subtitles right here. Bleeping, translation, burned-in subtitles and files over 90 MB live in the full tool.

Transcribe audio now →
What speaker identification actually gives you

It's also called diarization: not recognising who someone is, but noticing that this voice and that voice are two different people, and marking where each one is talking. The result is a transcript you can read like a script instead of a wall of text.

From recording to transcript here

Every stretch of speech is turned into a voice fingerprint — an ECAPA-TDNN embedding — and those fingerprints are clustered: voices that sit close together are treated as one person. Whichever cluster a line falls into decides its label, and the label is written into the TXT, the SRT and the VTT.

It runs by itself. There's no setting to switch on and nothing to tell it in advance, including how many people are in the room.

Where it stops — honestly

It can't know anyone's name. It has never heard your voice before, so it says Speaker 1, not "Sarah": read the first line each person says and replace the labels once in your text editor.

The transcript itself is unaffected either way: if the voices can't be told apart reliably, you simply get the normal transcript without labels.

Good with

Speaker labels come out of the same upload as everything else, so a recorded call can come back labelled, with subtitles, and with the swearing bleeped if you need to share it. See also transcribing an interview and a Zoom or Teams recording.

Only need the pages covering one speaker's part, not the whole transcript-turned-PDF? ScanReviver's split PDF pulls out just the page range you need.

Illustrated portrait of Sam RidderBy Sam Ridder — I build and run ScribeGrab on my own hardware, on my own. Who I am.

Frequently asked

Do I have to switch anything on?

No. If two or more distinct voices are heard, the labels are added by themselves, in the transcript and in both subtitle formats.

Will it use people's real names?

No, and no tool can without being told who's who beforehand. You get Speaker 1, Speaker 2 and so on; find the first line of each and replace the labels in your text editor.

How many speakers can it handle?

Two to eight. Above that the voice grouping isn't reliable enough to be worth showing, so it leaves the labels off rather than guessing wrong.

Why does my transcript have no labels?

Either it heard one voice, or it heard more than eight groups, which means the recording was too noisy or too overlapping to cluster confidently. The transcript itself is still complete.

Is speaker identification free?

Yes, like everything else here: no account, 10 recordings a day, no watermark, and your file is deleted right after processing.

Does ScribeGrab do audio transcription with speaker identification built in?

Yes — automatic speaker identification runs on every transcript with more than one voice, no separate speaker identification software or add-on needed. It's on by default and included in the free tier.

How does speaker detection transcription work here — is it a plugin or automatic?

Fully automatic. As soon as more than one voice is detected, the transcription is split into speaker turns with a label for each, so you get transcription with speaker labels out of the box rather than a plugin to install or a setting to find.

Is this an audio diarization tool, or a plain transcriber?

Both in one: it's a full transcription tool with diarization (speaker separation) included, not a separate step or a paid add-on.

What if it gets a speaker label wrong?

It happens most on crosstalk or two very similar-sounding voices. The labels are estimated by AI, so give overlapping stretches a quick check and swap the label directly in your text editor — the timestamps make it easy to find the exact spot.

What format do the speaker labels come in?

Plain "Speaker 1:", "Speaker 2:" and so on at the start of each turn, in both the transcript and the SRT/VTT subtitles — no special software needed to read them, just a text editor.

Can I do speaker diarization online without installing anything?

Yes, the whole process — upload, diarization, transcript — runs in the browser on our GPUs. Nothing to install, and the file is deleted right after processing.

How does speaker labeling work in a transcription?

The model listens for changes in voice as it transcribes and splits the text into turns the moment it detects a different speaker, tagging each turn Speaker 1, Speaker 2 and so on — no manual marking, no separate step.

How does speaker labeling transcription work, and what does a speaker label in transcription actually look like?

The model listens for changes in voice and splits the text into turns automatically, tagging each one Speaker 1, Speaker 2 and so on directly in the transcript — no manual marking needed.

How do I get a correct speaker label for every voice in the transcript?

It's automatic and usually accurate on clear audio with distinct voices; on crosstalk or very similar-sounding speakers, do a quick check and relabel any turn that's off — no re-upload needed for that.