Three kinds of missing text
Not every gap is a fault, so sort what you see first.
- Small omissions
- Fillers, repeated words and abandoned starts disappear, but the meaning is intact
- Missing sentences
- A clear thought is absent, often a quiet aside, a reply from across the room or a line spoken over music
- Missing stretches
- Minutes pass with no text at all, or the transcript ends long before the recording does
Small omissions are usually deliberate. Whisper-style models tend to write tidy text and drop many ums, repeats and restarts, which suits most readers but not every purpose; verbatim vs clean verbatim transcription covers when that matters. The other two kinds deserve investigation.
Common causes of skipped speech
- Quiet or distant speech. A speaker turned away from the microphone, a soft-spoken guest or a question from the audience may sit too close to the noise floor to be recognized.
- Music or effects over speech. When the bed is as loud as the voice, recognition becomes uncertain and short phrases drop out; noise in general is covered in speech recognition and background noise.
- Overlapping speakers. When two people talk at once, the recognizer usually follows the louder one and the other voice disappears.
- The wrong audio track. A file with several audio tracks, such as a screen recording with microphone and system sound on separate tracks, may present a track without the speech you expected.
- Stereo problems. A voice recorded on only one channel can end up much quieter after a downmix to mono, and two channels with opposite polarity can partly cancel each other.
- Long silences and pauses. Stretches with no speech correctly produce no text, but they can look like gaps; occasionally they also trigger invented text, a separate failure covered in Whisper hallucinations.
- Chunk boundaries. Long recordings are processed in pieces, and a word spoken right at a cut can occasionally be clipped.
- A language switch. A passage in a second language may be skipped or rendered in the wrong language.
Locate the gaps with timestamps
- Download the timestamped transcript or the SRT file rather than the plain text, so every line carries a time.
- Scan for jumps: places where one line ends and the next starts many seconds later, even though you know someone was speaking.
- Note each jump's start and end time, and check the last line of the transcript against the length of the recording.
- Play the original at each gap with headphones. Listen for the reason: a low voice, music, crosstalk, silence, or a different language.
- If a gap contains no audible speech at all, check the file itself: play it in another player and look at which audio tracks it has and which one is selected.
- If gaps line up with roughly half-minute intervals and each involves only a word or two, read about seam errors in audio chunking for speech recognition.
Fixes matched to each cause
Quiet speech. Raise the level of the quiet passage in an audio editor, or normalize the whole file, before re-running. Boosting helps when the speech is audible but low; it cannot recover words buried under noise.
Music under speech. Use a version without the music bed if you have one, such as the narration track from your edit. Otherwise, a speech-isolation tool may help, though it can introduce artifacts of its own.
Overlap. Recognition will not reliably separate two simultaneous voices. Transcribe the overlapping moments by ear, or, for future recordings, give each speaker a microphone and ask them to take turns.
Wrong track. Re-export the file with the speech as the only or main audio track; video files with multiple audio tracks shows how to check and choose. Stereo problems are fixed the same way, by exporting a mono mix from the good channel.
Re-run only the gap. Cut out the problem stretch with a little margin, apply any fix, and transcribe just that piece. This copies two minutes starting at 18 minutes without re-encoding:
ffmpeg -ss 00:18:00 -i interview.m4a -t 00:02:00 -c copy gap.m4a
To normalize the loudness of the extract at the same time, re-encode it to WAV instead:
ffmpeg -ss 00:18:00 -i interview.m4a -t 00:02:00 -af loudnorm gap.wav
Times in the new transcript start at zero, so add the extract's start time when you paste the recovered lines back.
A worked example: two minutes that vanished
Suppose a 50-minute research interview comes back with a transcript that jumps from 18:02 to 19:47. Listening at 18:00 reveals the interviewee leaned back from the table microphone while describing a sensitive episode, so the voice is clear but very soft. The researcher extracts 18:00 to 20:00, normalizes it and transcribes the two-minute file, which costs the 5-credit minimum (0.5¢). The recovered paragraph is pasted into the main transcript with the times shifted by 18 minutes, then checked against the audio once more because it will be quoted.
The second listen matters. Recovered passages come from the hardest audio in the file, so they deserve the closest proofreading; the routine in how to proofread an AI transcript applies.
Limits: what re-running cannot fix
- Speech masked by louder sound may not be recoverable by any recognizer. If you cannot make out the words yourself, a machine is unlikely to do better.
- Overlapping voices stay hard, and transcripts have no speaker labels to show who was talking.
- Heavy noise reduction or speech isolation can remove consonants along with the noise and create new errors.
- Boosting a quiet passage also boosts its hiss and room noise.
- Missing fillers are a style choice of the model, not a fault you can fix by re-running.
- A recording that cut out during capture has nothing to recover.
How mydubly processes speech and what it returns
In mydubly, the browser decodes the audio locally to 16 kHz mono and splits it into chunks of about 30 seconds, cutting each one at the quietest moment near the boundary so words are rarely split. Each chunk is compressed and sent over HTTPS, transcribed with Whisper on the server, and returned as timestamped segments. You get a plain transcript, a timestamped transcript, and SRT and VTT subtitles, optionally with a translated transcript.
A few properties matter here. If a file has several audio tracks, only one is used, so the speech should be the only or main track. Transcripts have no speaker labels. Leaving or refreshing the page during a job cancels it, and credits are returned if a job ends without a result. Transcripts cost 1 credit per minute with a 5-credit minimum per file, which keeps re-running a short gap cheap.
When another workflow fits better
- For legal, medical or published quotes where every word matters, have a person verify the gaps or use a human transcription service.
- For recordings full of crosstalk, such as lively group discussions, plan on manual transcription of the overlapping passages.
- For noisy field recordings, cleaning the audio in an editor first usually gains more than any amount of re-running.
Next step: check one gap today
Download the timestamped transcript of your file, find the largest jump, and listen to the original at that point. Once you know the cause, fix that stretch and run the extract through audio to text; video files can go through video to text directly. The timestamped transcript format page shows what the times look like.
Frequently asked questions
Why does my transcript leave out um, uh and repeated words?
Whisper-style models were trained largely on tidy subtitle-like text, so they tend to drop fillers, stutters and false starts even when they are clearly audible. That is useful for reading but not for true verbatim work. If you need every filler, add them during a listening pass.
Why does my transcript stop before the recording ends?
Check the end of the recording first: long silence, music or applause produces no text. If speech continues, the file may have been cut short during export or upload, or the selected audio track may end early. Play the file to the end in another player and confirm which audio track carries the speech.
Can I re-run only part of a recording?
Yes. Cut the missing section out with a few seconds of margin in an audio editor or with ffmpeg, then transcribe that extract on its own. Remember that the new timestamps start at zero, so add the extract's start time when you merge the lines back.
Will boosting the volume make the recognizer hear quiet speech?
Often, if the voice is audible but low. Normalizing raises speech toward a level the recognizer handles well. It does not help when the words are masked by noise or music, because boosting raises everything equally.
Why is one person missing from my interview transcript?
The likeliest causes are a separate audio track or channel for that person, or a voice much quieter than the other. Check whether the file has several tracks or a one-sided stereo channel, then export a single mixed track with both voices at similar levels.