No account. No daily cap.Others give you three files a day and call it unlimited. Files up to 60 minutes here, as often as you like.
Free, no sign-up, no watermark. An hour of audio takes about two minutes.
No file at hand? Paste a YouTube link below. Subtitles can be translated into 16 languages.
A pass doesn’t change the transcription and doesn’t make the transcript better — it adds what you do with the transcript afterwards. What’s on this page is the whole tool, not a trial version.
This is the whole tool, not a trial. Same model and the same transcript as with a pass — nothing here is switched off.
Handled by Polar, our payment provider — they take care of the payment and the VAT. No account here: your pass is stored in this browser.
Paste a YouTube link and it comes back in a second — and if that video has no subtitles we transcribe it anyway, which is where most tools give up. Interviews, lectures, podcasts, voice memos: files up to 60 minutes each, no account, no watermark.
Any common audio or video format. Transcription starts immediately: no account, no e-mail, no daily cap.
A speech-recognition model transcribes it on our GPU, detects the language automatically and labels who is speaking when there is more than one voice, usually far faster than real time.
Get the plain transcript as TXT and time-coded subtitles as SRT and VTT. Your upload is wiped automatically.
SRT/VTT captions for YouTube, Reels and TikTok: accessibility and reach, without paying per minute.
Turn recorded interviews, meetings and lectures into searchable text you can quote and skim, split by speaker so you can see who said what.
A full transcript per episode for show notes and SEO: every word your episode says, indexable.
Between 13 July and 21 August 2026, ScribeGrab logged 19,344 page views, 16,385 of them from real visitors rather than bots, arriving from 143 different countries — the US (421), India (303), Brazil (264) and Germany (153) sent the most, followed by the Philippines, Argentina, the Netherlands and France (112 each). Most of that traffic showed up as an ordinary browser visit (2,380) or moved between pages already on the site (1,484 internal), with only a handful opened from inside the Google or Android apps (9 combined). Across every tool here, 1,647 jobs actually ran in the same window: mostly audio to text (766) and video subtitles (695). Not everything dropped on the page finishes, though — 1,083 uploads were rejected before processing even started, and by far the most common reason was a file running past the 60-minute limit, hit 155 times in 40 days.
Yes: no account, no watermark, no per-minute charge or daily cap, and the free tool is not a trial — it doesn't expire and it runs the same model as everything else here. Ads keep the GPU running. There is one optional extra: a €5 month pass with meeting notes, real speaker names, Word export and higher limits. It buys you more to do with your transcript, not a better transcript.
It's processed on our own server and deleted the moment it's done; the results are wiped within 45 minutes. Nothing is stored, shared or used for training. See the Privacy Policy.
We use our speech model, one of the best open speech-recognition models. Clear speech transcribes very accurately; heavy accents, crosstalk, background noise or poor recordings are harder for any tool. Always give it a quick proofread.
Yes. When it hears more than one voice it splits the transcript into speaker turns automatically, up to eight speakers, with nothing to switch on. That's what makes an interview or a meeting recording readable instead of one long block of text. The labels are estimated by AI, so on crosstalk or very similar voices give them a quick check.
Yes. Tick “Also make a clean or subtitled version” before you upload and you get a second file back with every swear word beeped, muted or cut out — plus, if you want, the “um”s and the long silences. You see the full list with timestamps first and can leave any word in. More on bleeping swear words and removing filler words.
Around 100 languages, auto-detected (or pick one). Audio and video: MP3, WAV, M4A, FLAC, OGG, MP4, MOV, MKV, WEBM and more, up to 2 GB and 60 minutes.
Yes, upload the video directly and download the SRT or VTT file. Drop the SRT next to your video or upload it to YouTube as captions.
Yes. Instagram, TikTok, WhatsApp and most TVs ignore a separate subtitle file, so tick “Also make a clean or subtitled version” before you upload, pick one of three looks (Classic, Boxed or Big) and download an MP4 with the text burned into the picture. No editor to install, no watermark, and your original file stays untouched. Full guide: how to add subtitles to any video.
That's our sister tool: StemGrab separates vocals, drums, bass and instrumental, also free.
On a clear recording of one person speaking English the model gets roughly ninety-eight words in a hundred right — we measured 1.94% word errors on clean speech. Three things hurt most. Crosstalk, where two people talk over each other, because the model has to pick one. Distance from the microphone, since a phone on the far side of a meeting table picks up more room than voice. And unusual vocabulary: names, brands, jargon and abbreviations are guessed from context, so a surname it has never seen comes out as something that merely sounds like it. Strong accents are handled better than most people expect; background music is handled worse.
Speech recognition costs GPU time, and most free transcribers rent that time from a cloud service per minute of audio — which is exactly what the three-files-a-day and twenty-minute caps are there to protect. We own the hardware: the same graphics cards run our other tools in the Netherlands, so an extra recording costs electricity rather than an invoice, and the ads on the page cover it. That is why files can be up to ninety minutes and two gigabytes each, as many as you like, without an account or an e-mail address.