Every "um", every "uh", every four-second pause where nobody says anything: found automatically and cut out on the millisecond. Your recording gets shorter and tighter, and you never touched a timeline.
No account. 10 recordings a day. Others give you three files a day and call it unlimited. Files up to 60 minutes here, and it resets at midnight.
Plain transcript and subtitles right here. Bleeping, translation, burned-in subtitles and files over 90 MB live in the full tool.
Transcribe audio now →Unscripted speech is full of gaps. Thinking noises between sentences, the pause while someone finds a slide, the silence after a question. Individually they're a second or two; across an hour of talking they're minutes, and they're the difference between a recording people finish and one they close.
Cutting them by hand is the least interesting work in editing: find the gap, set two markers, ripple delete, repeat two hundred times.
Removing filler words manually means scrubbing the waveform second by second for every "um" — which is why most people give up after the first five minutes and just leave the rest in.
"Like" and "you know" are deliberately not on the list. They're usually part of the sentence, and cutting them leaves speech that reads fine on paper and sounds mangled out loud.
Nothing is sped up, pitched or time-stretched. Both the audio and, for video, the picture are cut at the same points, so voices sound exactly like themselves and lips stay on the words. What comes back is simply the same recording minus the parts where nothing was said.
Filler words and silences sit next to the swear-word options in the same panel, so one upload can hand back a recording that is shorter and bleeped. You still get the normal transcript, SRT and VTT out of the same job — see audio to text.
By Sam Ridder — I build and run ScribeGrab on my own hardware, on my own. Who I am.
That depends entirely on the speaker. A tightly scripted read loses almost nothing; a nervous unscripted interview can lose several minutes an hour. The result tells you how many pieces were removed.
It only cuts filler sounds and stretches where nobody is speaking, and it leaves a quarter-second of quiet on each side of a silence so the join doesn't sound abrupt.
No, on purpose. Those usually carry meaning or grammar, and removing them tends to leave sentences that sound broken.
Yes. Picture and sound are cut at the same points, so nothing goes out of sync. Video that is cut has to be re-encoded; beeping or muting alone copies the picture untouched.
Yes. 60 minutes and 1.5 GB per file, 10 recordings a day, no account and no watermark.
Yes — upload the recording and ScribeGrab strips the ums, uhs and dead air for free, no account, 10 recordings a day, up to 60 minutes per file.
Yes. Upload the MP4 or audio export of the call or screen recording and it works the same way — filler sounds and silences are cut, picture and sound stay in sync.
Yes — upload the recording and tick the filler-word option; every "um", "uh" and long pause is cut automatically, and the timestamps stay in sync so picture and sound don't drift.
Yes — upload the recording, tick the filler-word option, and every "um", "uh" and long pause is cut out automatically, with timestamps kept in sync. Works on any audio file, not just calls or podcasts.