How it's built

How to Split Long Audio for Speech Recognition Without Cutting Words in Half

Audio chunking for speech recognition means splitting a long recording into short windows, usually around 30 seconds, that a model such as Whisper can process one at a time. The splitting itself is trivial; where you cut is what matters. A boundary that lands inside a word produces garbled, missing or duplicated text on both sides, while a cut placed in a natural pause near each target length avoids most of those errors.

9 min read · Updated

Why speech models can't take a whole recording at once

Most modern recognizers are trained on short clips, and some have a hard input size. OpenAI's Whisper turns audio into a log-Mel spectrogram that covers exactly 30 seconds: anything shorter is padded with silence, and anything longer has to be cut before the model sees it. Models without a fixed window still run into the cost of attention, which grows with the square of the input length, so an hour of audio in a single pass would need vastly more memory than one short clip.

There are engineering reasons beyond the model as well. Short chunks can be processed in parallel, so a two-hour file does not need two hours of sequential compute. A chunk that fails can be retried on its own instead of restarting the whole job. Progress can be reported chunk by chunk rather than as a long wait followed by a single result.

So the real question is never whether to chunk a long recording, only how. That decision quietly sets the quality of every sentence near a boundary, and it is made before recognition even starts.

Fixed-length windows and what goes wrong at the seams

The simplest strategy cuts every 30 seconds by the clock. It is easy to implement, gives perfectly predictable chunk counts and ignores the content completely. That last property is the problem.

Speech is continuous. At an arbitrary instant there is a fair chance someone is mid-word and a good chance they are mid-sentence. When a cut lands inside a word, each side receives a fragment of sound the model was never trained to hear in isolation. Typical symptoms:

  • The end of the first chunk finishes with a misheard partial word, or a plausible whole word the speaker never said.
  • The start of the next chunk begins with a fragment that is dropped or turned into a different word.
  • Punctuation and capitalization break, because every chunk assumes a sentence begins at its first sample.
  • A fragment followed by padding can push the model to invent filler text, one of the failure modes covered in our article on Whisper hallucinations.

With 30-second chunks, a one-hour recording has about 120 seams. Even if only some of them land badly, the result is a steady trickle of errors in places a reader cannot predict and a proofreader has no reason to look.

Overlapping windows and the stitching problem

A common fix is to overlap the windows, for example 30-second chunks that start every 25 seconds, so each seam is covered twice. A word that is split at one boundary appears whole in the neighboring chunk.

Overlap moves the difficulty rather than removing it. You now hold two transcripts of the same five seconds that rarely agree word for word, and you must decide which words survive. Pipelines resolve this by matching timestamps, by aligning the two texts with an edit-distance search, or by keeping whichever half of the overlap sits nearer each chunk's center, where the model had the most context. These methods work most of the time and occasionally duplicate or drop a phrase. Overlap also costs extra compute and, in a cloud pipeline, extra upload.

Whisper's reference implementation takes a third route for long files. It transcribes one window, slides the next window to start at the last complete segment's timestamp, and passes the previous text in as a prompt. That preserves context but is inherently sequential, and conditioning on earlier output is one of the ways a repetition loop can start and then spread through the rest of the file.

Pause-aligned chunking: cut where the speaker stops

For file-based transcription, a robust approach is to keep chunks non-overlapping but move each cut into a natural pause. People breathe, finish sentences and hesitate, and 30 seconds of real speech almost always contains several gaps of a few hundred milliseconds. If the boundary falls in one of those gaps, both chunks receive complete words and usually complete sentences, and there is nothing to stitch afterwards.

Finding a pause requires some measure of where the audio is quiet. Two families of tools exist:

  • Energy measures compute the loudness of short frames, commonly as root mean square amplitude, and look for low values. They are fast, need no model and behave the same in every language.
  • Voice activity detection models classify each frame as speech or non-speech. They cope better with noisy rooms and music, at the price of running a second model.

For placing cuts, as opposed to discarding non-speech, plain energy is usually sufficient. You are not asking whether a frame contains speech, only which frame within a short range is the least harmful place to cut, and the quietest frame almost always sits in a gap between words.

How mydubly cuts a long recording into windows

mydubly uses pause-aligned windows of about 30 seconds, planned in the browser before any audio leaves the device. As implemented, the procedure is:

  1. A WebAssembly build of ffmpeg decodes the file's audio track once to 16 kHz mono 16-bit PCM. Whisper expects 16 kHz input, so no further resampling is needed later.
  2. The browser measures root mean square loudness for every 50 ms frame, which at 16 kHz is 800 samples per frame.
  3. Starting at zero, it sets a target end 30 seconds ahead, then scans the 6 seconds before that target, 120 frames, for the quietest one. The cut goes in the middle of that frame. If several frames are equally quiet, the latest wins, so windows stay as long as possible.
  4. The next window starts exactly where the previous one ended, so there are no gaps and no overlap to reconcile.
  5. If fewer than 6 seconds would remain after a window, that tail is folded into the final window rather than being sent as a tiny chunk of its own, so the last window can run a little past 30 seconds.

Every window except the last therefore lasts between 24 and 30 seconds. Each one is encoded to Ogg Opus at 32 kb/s with the browser's built-in WebCodecs encoder, roughly 120 KB per 30 seconds, and several upload in parallel; the article on reducing upload size covers that step. On the server, recognition runs on OpenAI's Whisper, with large-v3-turbo as the default model, and returns text segments with start and end times for each chunk.

Worked example: a 40-minute lecture

The first target is 30.0 s. Between 24.0 and 30.0 s, the quietest 50 ms frame runs from 27.40 to 27.45 s, where the lecturer pauses after a sentence, so window one is 0 to 27.425 s. Window two targets 57.425 s and finds a breath at 55.1 s. Continuing this way, the 2,400-second file becomes at least 80 windows, in practice a few more because cuts land early. Near the end, a remainder shorter than 6 seconds joins the last window. As a transcript job, the whole lecture costs 40 credits, which is 4 cents.

Why the cut point matters beyond recognition

In a translation pipeline the chunk is more than a recognition unit. In mydubly, each chunk's segments are translated, voiced and timed together, and each chunk of generated speech is built to be exactly as long as its window so the pieces join without gaps. A boundary inside a sentence would split that sentence across two translation requests and two pieces of synthesized speech, which is worse than a split in a transcript because the viewer hears it.

Cutting at pauses means a sentence nearly always lives inside a single chunk. The translation engine receives a complete thought and the voice receives a complete line to read. It also means each chunk's timestamps can be shifted cleanly into the full timeline, the process explained in subtitle timestamp alignment.

When pause-aligned chunking helps most

  • Lectures, talks and narration, where speakers pause regularly between long sentences.
  • Interviews and podcasts with clean turn-taking, because the gap between two speakers is an ideal cut.
  • Very long files, since seam errors scale with the number of seams and an hour already has more than a hundred.
  • Pipelines that translate or dub, where one split sentence damages several downstream stages at once.
  • Parallel processing, because non-overlapping chunks are fully independent and need no merge step at the end.

Trade-offs and cases where boundaries still cause errors

No chunking strategy is free. These are the limits of the approach above, roughly in the order you would notice them:

  • Continuous speech with no pause in the 6-second search range. A fast talker, a chant or a song may never stop. The cut still lands on the quietest frame, usually a gap between syllables, but it can clip a word.
  • Music or noise under the speech. Energy measures loudness, not speech, so under a constant music bed the quietest frame may still fall inside a word. A voice activity model would do better here.
  • Lost context at the start of each chunk. Each chunk is recognized independently, so the model does not know the previous sentence. Expect occasional inconsistency in how a name or rare term is spelled from one chunk to the next.
  • Translation context also stops at the chunk edge. A pronoun whose referent sat in the previous chunk can come out with the wrong gender or formality, a problem discussed in context in machine translation.
  • Variable chunk lengths. Pause alignment gives up perfectly predictable chunk counts, which matters little to users but does matter to systems that schedule work per chunk.

For practical advice on long recordings, including what to do with files over the 2-hour limit, see transcribing long audio files.

Spotting seam errors in a finished transcript

If you suspect chunk boundaries are hurting a transcript, from any tool, this check takes a few minutes:

  1. Export a timestamped transcript or an SRT file so you can see where each segment begins.
  2. Look near each half-minute mark for segments that start or end abruptly, and read the last words before and the first words after.
  3. Listen to those moments at normal speed. A doubled word, a missing word or a nonsense word right at a boundary points to a seam problem rather than general recognition error.
  4. If the errors cluster in a noisy or musical stretch, cleaner source audio will help more than any chunking change; our guide to improving transcription accuracy lists what to fix.

Run your own recording through pause-aligned chunking

The easiest way to judge the effect is to transcribe a long recording with mydubly's video to text tool and spot-check a few segment starts around the 30-second marks. Audio-only files go through audio to text the same way. Transcripts cost 1 credit per minute with a 5-credit minimum per file, so a one-hour recording uses 60 credits, leaving 40 of the 100 free credits a new account receives when signing up with Google.

Frequently asked questions

Why is 30 seconds the usual chunk length for Whisper?

Whisper's encoder was trained on 30-second log-Mel spectrograms, so every input is padded or cut to that length. Chunks close to 30 seconds waste the least of that window. Much shorter chunks create more seams and more padding, giving the model less context and more silence to misread.

Should audio chunks for transcription overlap?

Overlap protects words from being split, but you then need logic to merge two slightly different transcripts of the shared section, which can duplicate or drop phrases. If cuts can be placed in pauses, non-overlapping chunks are simpler and skip the merge entirely. Overlap remains useful for audio that never pauses, such as continuous singing.

Is voice activity detection better than loudness for finding cut points?

For deciding which frames contain speech, a voice activity model is more reliable, especially with music or crowd noise. For picking the quietest moment within a few seconds, frame loudness is cheap and usually finds a real gap between words. A sensible design uses loudness for cut placement and keeps a voice activity model for skipping non-speech.

Does chunking shift the timestamps in the final transcript?

Each chunk's segment times start at zero, so they must be offset by the chunk's start time in the full recording. With non-overlapping, pause-aligned chunks those start times are exact, so the offset adds no error. The full process is described in subtitle timestamp alignment.

Can I choose where mydubly splits my file?

No. Cut points are chosen automatically from the audio's loudness, roughly every 30 seconds. If you need a particular split, for example to bring a recording under the 2-hour limit, trim it at a pause in an editor before uploading.