The punctuation problems people actually see
Punctuation errors in transcripts are rarely random. They cluster into a few recognisable shapes, and naming the shape tells you which fix to reach for.
- Run-on sentences. A speaker strings ideas together with "and", "so" and "but", and the transcript turns three minutes of talk into a handful of enormous sentences.
- Chopped sentences. A thoughtful speaker pauses mid-thought, and the transcript puts a full stop in the pause, leaving fragments such as "We decided to. Move the launch."
- Comma storms or comma droughts. Some passages get a comma after every few words, others none at all, depending on how the speaker phrased them.
- Questions without question marks. Questions asked with falling intonation, or statements that rise at the end, get the wrong closing mark.
- Capitalisation slips. Product names and job titles appear in lower case, ordinary words get capitals at false sentence starts, and "i" occasionally survives.
- Inconsistent style across a long file. One section uses dashes for interruptions, another uses ellipses, another uses nothing.
How speech recognition decides where punctuation goes
Older speech recognition systems produced a stream of lower-case words and added punctuation in a separate step, often with a model trained on written text. Modern end-to-end models such as Whisper write punctuation and capitals directly, as tokens alongside the words, because they learned from large amounts of audio paired with punctuated transcripts. The walkthrough of how speech to text works covers that decoding process.
Either way, the model is inferring something the speaker never said. It weighs the words themselves (a clause that sounds finished), the acoustics (a falling pitch, a pause, a breath) and the habits of the transcripts it was trained on. Training transcripts vary in style: some are tidy subtitles, some are rough captions, some are verbatim. That variety is one reason the same model can punctuate two similar passages differently.
Why pauses, pace and speaking style change sentence breaks
Pauses are the strongest acoustic hint, and they are unreliable. People pause to breathe, to find a word or for emphasis, not only at the end of a sentence. Conversely, fluent speakers often run from one sentence into the next without any gap at all.
- Fast, continuous speech hides boundaries, so you get run-ons.
- Slow, deliberate speech with long thinking pauses produces fragments.
- Lists spoken in a flat rhythm often lose their commas.
- Interrupted or overlapping speech produces half-sentences that end abruptly or join two speakers together.
- Reading from a script usually punctuates well, because written sentences are spoken with clear endings.
Recording quality matters too. When noise or echo blurs the end of a word, the model has less evidence about where a phrase finishes. Fixing the sound does more for punctuation than any setting in a transcription tool.
Chunk boundaries and sentence breaks
Long recordings are not transcribed in one pass. They are split into windows, commonly around 30 seconds, because that is the span models such as Whisper were trained on. If a cut lands in the middle of a sentence, the first window ends on half a thought and the next starts with the other half, and the model may close the first half with a full stop and capitalise the second.
Systems reduce this by cutting where the audio is quietest, which usually means a pause between phrases rather than the middle of a word. Even then, a pause is not always a sentence end, so an occasional false break near a cut is a normal property of chunked transcription. The article on audio chunking for speech recognition explains the trade-offs between fixed and pause-aligned windows.
Diagnose before you start editing
- Read the first two or three minutes and note which patterns from the list above appear. Most files have one or two dominant problems.
- Check whether the errors are spread evenly or concentrated in certain passages. Concentrated problems usually track a speaker, a noisy stretch or a change in speaking style.
- Use a timestamped transcript to listen to a few suspicious sentence breaks. If the speaker really did pause there, the model behaved reasonably and you are choosing an editorial style, not correcting an error.
- Decide the target style. A searchable archive needs far less polish than a published interview or a subtitle file.
- Only then pick the fixes below, starting with the cheapest.
Fixing punctuation by hand and with find-and-replace
Mechanical problems are best handled with find-and-replace in any text editor. Do them on a copy, one pass at a time, and skim the result after each pass. If your editor supports regular expressions, a few patterns go a long way:
- Double spaces
- Replace two spaces with one, and repeat until nothing is found
- Space before punctuation
- Find a space followed by a comma or full stop and remove the space
- Lower-case "i"
- Search for " i " and " i'm " as whole words and capitalise them
- Product and people names
- Replace each misspelt or lower-case name with the correct form, using whole-word matching
- Spoken fillers that became sentences
- Search for lines or sentences consisting only of "So." or "Okay." and decide whether to keep them
Some editors can also change case in a replacement, which helps with capitals after full stops; check your editor's documentation because the syntax differs. Everything else is judgement and is faster by hand: splitting run-ons where the speaker changes topic, joining fragments across a thinking pause, and choosing between a comma, a dash and a full stop. Reading aloud while listening at a slightly faster playback speed is the quickest way to hear where a sentence really ends. The guide to proofreading an AI transcript covers a full review pass, and formatting a transcript covers paragraphs and layout once the sentences are right.
A worked example: a podcast episode with run-on answers
Suppose a host transcribes a 50-minute interview with a guest who speaks quickly and links every idea with "and so". The transcript reads well in the host's questions but turns the guest's answers into sentences of 80 words or more. The host first runs the mechanical passes: double spaces, the guest's company name, and the lower-case product name. Then, listening at 1.25 times speed, they split each long answer where the guest changes subject, which is usually just before "and so" or "but the thing is". The edit takes about as long as one listen-through, and the meaning of every answer is checked against the audio rather than guessed from the text.
What punctuation does to subtitles and translation
Punctuation is not cosmetic once a transcript feeds other outputs. Machine translation works sentence by sentence, so a run-on gives the translation engine a tangle to resolve and a false full stop splits one idea into two unrelated sentences. Fixing sentence boundaries in the source transcript before translating anything by hand will usually improve the result more than editing the translation afterwards.
Subtitles have their own conventions: short lines, a full stop where a thought ends, and no trailing comma that leaves a cue hanging. For synthetic voices, punctuation steers phrasing and intonation, so a missing question mark or a misplaced full stop changes how a line is spoken.
Mistakes and limits of punctuation cleanup
- Correcting style instead of errors. A pause the speaker really made can reasonably be a full stop; rewriting it may not improve anything.
- Automating judgement calls. A global replace of "and so" with ". So" will be wrong in many places.
- Losing the link to the audio. Heavy restructuring of a timestamped transcript makes it harder to find the original passage later; keep the raw file.
- Expecting punctuation to fix wrong words. If the recognised words are wrong, commas will not rescue the sentence; see why transcription can miss words for that problem.
- Over-punctuating spoken speech for publication. Readers accept shorter, plainer sentences than speakers produce, but changing a quote's meaning is never acceptable.
How mydubly handles punctuation
mydubly transcribes with Whisper large-v3-turbo in its default deployment, so punctuation and capitals come from the model's predictions, not from a separate rule set, and the same patterns described above apply. The audio is split into chunks of roughly 30 seconds, with each cut placed at the quietest 50 ms moment in the last 6 seconds of the window, so cuts tend to fall in pauses. Each chunk is transcribed without conditioning on the previous chunk's text, a setting commonly used to reduce the risk of repetition loops.
You download the result as plain text, as a timestamped transcript with [m:ss] labels, and as SRT and VTT subtitles with one single-line cue per recognised segment. mydubly has no built-in transcript editor, so punctuation fixes happen in your own text or subtitle editor. If you also need a translation, punctuation in the downloaded translated transcript comes from the translation engine, so check it with the same eye.
Next step: transcribe a test file and set a house style
Run a short, representative recording through audio to text, read the result against the audio for five minutes and write down which punctuation patterns appear. A transcript costs 1 credit per minute with a 5-credit minimum, so a 12-minute test costs 12 credits (1.2¢). Turn your fixes into a short checklist of find-and-replace passes, and the next transcript will take a fraction of the time to clean up.
Frequently asked questions
Why does my transcript have almost no full stops?
Fast, continuous speech gives the model few acoustic clues about sentence ends, so it tends to join clauses into long run-ons. Speakers who link ideas with 'and' or 'so' make this worse. Split the text where the speaker changes topic, and listen to the audio wherever the split could change the meaning.
Why does the transcript break sentences in the middle?
Models treat long pauses as likely sentence ends, so a speaker who stops to think can end up with fragments. A cut between two transcription chunks can occasionally have the same effect. Joining the fragments is quick once you have confirmed in the audio that the speaker carried on with the same thought.
Can I make a tool punctuate more consistently?
Mostly by improving the input: clean audio, one speaker at a time and natural pacing give the model clearer cues. Beyond that, consistency comes from editing to a house style. A short list of find-and-replace passes applied to every transcript does more than switching between similar recognition models.
Should I fix punctuation before or after translating?
Before, if you are going to translate or post-edit the text yourself. Translation works sentence by sentence, so clean sentence boundaries in the source lead to cleaner translations. If a tool has already produced both transcripts, fix the source first, then check the matching translated sentences.
Why are questions missing question marks?
Many questions in speech are asked with falling intonation, or as statements with a rising tone, so the model has to guess from the words alone. Search the transcript for question words such as 'what', 'how' and 'why' at the start of sentences and check how each ends.
Does punctuation affect AI dubbing?
Yes. Speech synthesis uses punctuation to decide phrasing and intonation, so a missing full stop can make two ideas run together and a stray comma can add an odd pause. Clean sentence boundaries in the transcript help a translated voice sound more natural.