No account. No daily cap.Others give you three files a day and call it unlimited. Files up to 60 minutes here, as often as you like.
Upload a recording, get the words back as a text file and as subtitles, and the audio is deleted the second the transcript exists.
You drop in a recording, a graphics card in my house runs it through our speech model, and you get back a plain-text transcript plus time-coded SRT and VTT subtitles. No account, no watermark, no daily cap, no per-minute meter ticking along in the corner.
The "my house" part is literal. I am one person with one machine, and I also built StemGrab and PicReviver, which run on the same two cards. When all three are busy your file waits in a queue behind somebody's song. That is the honest reason a transcript is sometimes instant and sometimes takes a minute.
The hard numbers, so you know before you upload. Every value here is what the server actually enforces.
| Free | With a pass | |
|---|---|---|
| Price | Free. No account, no sign-up, no card. | €5 for 30 days, or €3 for 24 hours |
| Maximum length | 60 minutes per file | 4 hours per file |
| Maximum file size | 1.5 GB | 1.5 GB |
| Files at once | 2 | 5 |
| Queue | Normal | Priority when the cards are busy |
| Results kept for | 45 minutes | 24 hours |
| Speaker labels | Added automatically on both tiers — "Speaker 1:", "Speaker 2:" in the transcript and subtitles | |
| Renaming the speakers | No | Yes — replace "Speaker 1" with a real name throughout |
| Word (.docx) export | No | Yes |
| Meeting minutes | No | Yes |
| Translate into | One language at a time | Up to 5 at once |
| Transcribing per day | No cap | No cap |
| Text-to-speech per day | 10 minutes | 60 minutes |
| Text tools per day | 25,000 characters | 250,000 characters |
| Input formats | 25 audio and video formats: MP3, WAV, FLAC, M4A, AAC, OGG, OPUS, WMA, AIFF, AIF, MP4, MOV, MKV, WEBM, AVI, M4V, MPEG, MPG, WMV, 3GP, AMR, M4B, CAF, TS, FLV | |
| Output formats | TXT, SRT, VTT | TXT, SRT, VTT, DOCX |
| Transcription languages | About 100, detected automatically | |
| Translation languages | 16: English, Spanish, French, German, Dutch, Portuguese, Italian, Polish, Turkish, Russian, Arabic, Hindi, Indonesian, Japanese, Korean and Simplified Chinese. Timestamps are kept. | |
| Where it runs | On our own graphics cards in the Netherlands. Your audio is never sent to OpenAI, Google, AWS or any other transcription service. | |
| Your upload | Deleted the moment processing finishes, success or failure — on both tiers. | |
| Watermark | None, on any output, on either tier. | |
The free tier is not a trial and does not run out. What the pass buys is length, speed and the extra export formats — the transcription itself is the same model either way.
Otter, Rev, Sonix and the rest pay a cloud provider for every minute of audio they touch, so they have to charge you by the minute or cut you off after a few files. I do not have that bill. The cards are bought and paid for; your hour of audio costs me some electricity and nothing else.
The ads on the page cover that electricity and the domain name. There is also an optional pass — €5 for 30 days, €3 for 24 hours — which raises the limit from 60 minutes to 4 hours, adds speaker labels, Word export and meeting minutes, and puts you ahead in the queue when the cards are busy. The exact difference is in the table above. What it does not do is take anything away from the free tier: the SRT was always going to be free, it is the same file with timestamps in it, and no free limit was lowered to make the pass look better.
Transcription only works if you hand me your recordings, so the rules are short and I do not bend them:
The full details are in the Privacy Policy.
Clear speech — interviews, lectures, podcasts, voice memos, screen recordings — comes back very accurate, in around 100 languages that it detects by itself. Three things beat it, and they beat every other system too: a room full of background noise, people talking over each other, and a very strong accent on a bad microphone. I wrote an honest guide about how accurate automatic transcription really is rather than quote a percentage at you. Music does not transcribe at all; for splitting songs into stems, use StemGrab.
Found a bug, got a strange result, or just want to say something? The contact form lands in my inbox, not a ticket queue.
About the byline. I write as Sam Ridder. It is a pen name: transcription is a thing people hand private recordings to, and I would rather they judged the setup I have described here than a name they can look up. Every factual claim on this page is accurate.