Speech recognition & transcription

Transcribing long audio files: chunking, boundaries and the 2-hour limit

To transcribe long audio files, speech recognition systems split the recording into short windows, about 30 seconds each for Whisper-based tools, transcribe every window and stitch the timestamped results back together. Where those cuts fall matters, because a cut through the middle of a word can drop or duplicate it. In mydubly, files of up to 2 hours are split automatically at pauses; longer recordings should be divided into parts at a natural break before you upload them.

7 min read · Updated

Why length is a problem for speech models

A speech model never hears an hour of audio at once. Whisper was trained on 30-second windows: its encoder has a fixed number of positions, 1,500 after downsampling, and its decoder writes at most a few hundred tokens per window. Transformer attention also gets more expensive as the input grows, with cost rising roughly with the square of the sequence length, so simply feeding in longer inputs isn't practical. Every long-form transcription is therefore a sequence of short transcriptions joined together.

That makes long recordings an engineering problem as much as a recognition problem. How the audio is cut, whether chunks are processed one after another or in parallel, and how their timestamps are combined all affect the final transcript. The underlying model may be identical in two tools that produce noticeably different results on the same two-hour file.

Chunking strategies compared

Fixed windows
Cut every 30 seconds regardless of content. Simple and predictable, but cuts often fall mid-word
Overlapping windows
Fixed windows that share a few seconds with their neighbors; duplicated text in the overlap must be detected and merged
Sequential, timestamp-driven
Whisper's reference approach: transcribe a window, then start the next one where the last complete segment ended. Clean boundaries, but strictly one window at a time
Voice-activity-based
A voice activity detector finds speech regions and groups them into chunks at silences; skipping non-speech also reduces invented text
Pause-aligned cutting
Aim for a target length, then move each cut to the quietest moment near it, so chunks stay close to the model's window and boundaries fall in natural pauses

Each choice trades speed against context. The sequential approach can feed the previous window's text into the next as a prompt, which helps keep spellings and style consistent, but an error or a repetition loop can then propagate from window to window, and nothing can run in parallel. Independent chunks can be transcribed simultaneously, which is much faster for long files, but each chunk starts without knowing what came before. Audio chunking for speech recognition goes deeper into the engineering of these choices.

What goes wrong at chunk boundaries

  • Split words: half a word ends one chunk and half begins the next, and the model may guess at both fragments or drop them.
  • Broken sentences: a sentence that starts in one chunk and finishes in the next can gain a stray full stop or capital letter at the seam.
  • Duplicates: overlapping schemes can repeat a phrase if the merge step misjudges the overlap.
  • Invented filler: a final chunk that is only a few seconds long and mostly silent invites hallucinated text.
  • Timestamp offsets: each chunk's times are relative to its own start, so every chunk's offset must be added correctly, or subtitles drift. Subtitle timestamp alignment explains that step.
A cut through a word

Suppose a fixed cut lands at exactly 30.00 seconds while the speaker is saying "infrastructure". Chunk one ends "we need to rebuild the infra" and chunk two begins "structure team's budget". The stitched transcript reads "rebuild the infra structure team's budget", or one chunk drops its fragment entirely. Moving the cut back to the short pause before "we need" keeps the whole phrase inside chunk two and avoids the error.

Keeping a long transcript consistent

Long files expose problems that short ones rarely show. A guest's surname might be spelled one way at minute 5 and another way at minute 85. A product name might alternate between two forms. Recording conditions might change after a break, and the quality of the transcript changes with them. When chunks are recognized independently, nothing forces one window to agree with another.

  1. Before reviewing, write down the names, places and technical terms you expect to hear.
  2. Search the transcript for each one, including likely misspellings, using just the first few letters to catch variants.
  3. Replace the wrong forms throughout the document in one pass.
  4. Listen closely wherever the audio changes character: after breaks, during music, when a new speaker joins.
  5. Check the last minutes and any long silences for lines that nobody actually said.

How mydubly handles long files

In mydubly, the heavy lifting starts in your browser. A WebAssembly build of ffmpeg decodes the file's audio once into 16 kHz mono PCM. The source file is mounted rather than copied into WebAssembly memory, so even a multi-gigabyte video works; the decoded audio for two hours is about 230 MB. The audio is then split into windows of about 30 seconds, and each cut is moved to the quietest 50 ms frame within the last six seconds before the 30-second mark. A short leftover tail is folded into the final window instead of becoming a tiny chunk of its own.

Each window is compressed to Ogg Opus at 32 kb/s, about 120 KB per 30 seconds, and several chunks upload in parallel. A two-hour file becomes roughly 240 chunks totaling around 29 MB, far less than the original video. Whisper large-v3-turbo returns timestamped segments for each chunk, and those are placed on the full timeline to produce one transcript and one set of SRT and VTT files. Keep the tab open while the audio is extracted and uploaded, since that work happens in the browser.

Files longer than 2 hours: how to split them

mydubly accepts up to 2 hours per file, so a three-hour recording needs to be divided first. Cut at a pause, not at a round-number timestamp, for the same reason chunk boundaries should fall in silences.

  1. Find a natural break near the point you want to split, such as a pause between questions or a session break, and note its timestamp.
  2. Split the file in an audio or video editor, or with ffmpeg. For example, "ffmpeg -i recording.m4a -t 01:35:00 -c copy part1.m4a" writes the first 95 minutes, and "ffmpeg -i recording.m4a -ss 01:35:00 -c copy part2.m4a" writes the rest. Stream copying avoids re-encoding, though the cut may shift slightly to the nearest audio frame.
  3. Name the parts in order so they are easy to reassemble.
  4. Transcribe each part separately.
  5. If you need a single timeline, shift the second part's timestamps by the split time; most subtitle editors can offset an SRT or VTT file in one step.
  6. Join the transcripts and review the seam between the parts.
Splitting a 3-hour-10-minute recording

Suppose an oral history session runs 3 hours 10 minutes. Splitting it at a pause at 1:35:00 gives two 95-minute files. Transcribing both costs 95 + 95 = 190 credits, or 19 cents, and the result is two timestamped transcripts that join into one after a timing shift.

Benefits of chunked transcription

  • Length stops mattering to the model, since every step processes the same small window.
  • Chunks can be transcribed at the same time, so total time grows far more slowly than a strictly sequential pass would.
  • Memory and upload sizes stay bounded, which keeps long recordings practical on ordinary connections.
  • Segments map naturally onto subtitle cues.
  • With pause-aligned cuts, most boundaries fall in silences and cause no errors at all.

Limitations to plan for

  • Independent chunks don't share context, so a name the model got right early on may be misspelled later.
  • Speakers who talk quickly without pausing leave few quiet moments, so even a smart cut may land inside speech.
  • Long silences and music-only passages invite invented text in any Whisper-based system.
  • The 2-hour cap per file adds a splitting step for marathon recordings.
  • Review time scales with length; no chunking strategy removes the need to proofread a long transcript.
  • Transcripts don't label speakers, so attributing lines in a two-hour panel is manual work.

Next step: transcribe your long recording

Open the file in audio to text, or video to text if it's a video, and let the browser handle extraction and chunking. Transcripts cost 1 credit per minute, so a full two-hour file is 120 credits, 12 cents. For field-specific advice, see research interview transcription or lecture transcription, and if you also need the long recording translated, read translating long videos.

Frequently asked questions

What is the longest file mydubly can transcribe?

Up to 2 hours per file, for both audio and video. Longer recordings need to be split into parts first, ideally at a natural pause, and each part is transcribed and billed separately at 1 credit per minute.

Why does my transcript have errors every 30 seconds?

That pattern usually points to fixed-length chunking cutting through words. Tools that move cuts to pauses or use voice activity detection avoid most of it. If you split files yourself, cut in silences rather than at round-number timestamps.

Does a longer file make transcription less accurate?

Not inherently, because the model only ever sees one short window at a time. Long files simply show more of the problems that occur occasionally: inconsistent spellings of names across the recording, invented text during long silences, and changes in recording conditions partway through.

Should I trim silence before transcribing a long recording?

Trimming long dead air at the start, the end and during breaks is worthwhile, because it lowers the chance of invented text and saves a little cost. There's no need to remove the natural pauses between sentences, which help a chunking system find good cut points.

How do I join transcripts from two parts of a split recording?

Paste the second transcript after the first. If you need one continuous timeline, shift the second part's timestamps by the time at which you split; most subtitle editors have a shift-timing function that does this for SRT or VTT files in one step.