Why delivery matters as much as equipment
A good microphone in a quiet room captures your voice faithfully, including every mumble, overlap and trailing sentence. Speech recognition can only transcribe what it can make out, and a transcript reader can only follow sentences that were finished. Many of the errors people blame on transcription software trace back to delivery: words lost when a speaker turned away, two people talking at once, or a name nobody could spell.
The good news is that clear delivery does not mean sounding like a newsreader. It means a handful of habits that also make you easier to listen to. The underlying factors that make speech understandable, to people and machines, are explained in speech intelligibility.
Keep a consistent distance from the microphone
Distance changes everything about how a voice is recorded: level, tone and the amount of room echo. A speaker who drifts from 10 cm to 50 cm away becomes quieter, thinner and more reverberant, and the words at the far end of that drift are the ones most likely to be misrecognized.
- Find a comfortable position and stay in it. For a desk microphone, a hand's width away is a common starting point.
- Turn your head with the microphone if you look at notes or another person, rather than turning away from it.
- Avoid leaning back in a swivel chair or pacing away from a fixed microphone mid-sentence.
- On a phone recording, keep the phone in the same place on the table rather than picking it up and putting it down.
- If you use a headset or clip-on microphone, distance takes care of itself, which is one reason they suit long sessions.
Pace and articulation
Speak at your natural conversational pace, or slightly slower. Very fast speech packs more words into each second and blurs boundaries between them; very slow, over-careful speech sounds stilted and does not improve recognition much. Aim for steady rather than slow.
Articulation is about finishing words. The most common problems are dropped endings, where "going to" becomes a mumble or the last syllable of a sentence falls away as the voice drops, and words swallowed while laughing or eating. A useful habit is to keep your voice up to the end of each sentence instead of trailing off.
Avoid exaggerating consonants or shouting for clarity. Over-enunciation adds plosive pops and harsh s sounds, and raising your voice changes its tone and can clip the recording. If a recording you already have is simply fast, transcribing fast speech covers what to do with it.
Pause between thoughts
Short pauses at the ends of sentences and between ideas do several jobs. They give listeners time to absorb a point, they give a recognizer clear boundaries between phrases, and they often become the natural breaks between subtitle lines. Pauses are also where editors cut, so a pause before each new answer makes editing quicker.
The pause does not need to be long. A beat of half a second to a second between thoughts is enough. What hurts is the opposite: running sentences together with "and, so, and then" for minutes at a time, which produces transcripts with sprawling, hard-to-punctuate sentences.
Filler words such as "um" and "you know" are normal. You do not need to eliminate them, but replacing them with a silent pause when you can makes both the audio and the transcript cleaner.
Names, numbers, acronyms and jargon
Speech recognition systems recognize common words well and unfamiliar ones less well. Personal names, company and product names, place names and specialist terms are where most errors cluster. A few habits help the person who checks the transcript:
- Spell unusual names once, early: "I'm speaking with Siobhan Ní Bhriain, that's S-I-O-B-H-A-N." The recognizer may not use the spelling, but the proofreader will have it on the record.
- Use full names before short forms: "the Food and Drug Administration, the FDA" rather than only the acronym.
- Say numbers unambiguously. "Fifteen, one five" removes confusion with "fifty"; for dates, include the month name.
- Give context for jargon the first time it appears. Words that appear in a meaningful sentence are easier to correct later than isolated terms.
- Keep a list of names and terms as you record. It speeds up proofreading considerably; how to proofread an AI transcript shows how to use one.
One voice at a time
Overlapping speech is one of the hardest things for any transcription system, and for human transcribers too. When two people talk at once, the recognizer typically picks up one voice, mixes words from both, or drops the passage. Interjections like "yeah" and "right" during someone else's answer are less damaging, but they still add noise to the transcript.
- Agree on turn-taking before you start, especially in panels and group recordings.
- Hosts can signal turns by name: "Maria, what's your view?" This also helps readers identify who is speaking when the transcript has no speaker labels.
- Keep acknowledgments silent where possible: nod instead of saying "mm-hm".
- On remote calls, the delay makes accidental overlaps more likely; leave an extra beat before replying. Setup for those sessions is covered in recording a remote interview.
Briefing guests before you record
Guests rarely think about any of this unless someone tells them. A short, friendly briefing before the recording usually takes two minutes:
- Show them where to sit and how far from the microphone, and ask them to stay roughly there.
- Mention that the recording will be transcribed, so finishing sentences and spelling names helps.
- Ask them to silence their phone and avoid tapping the table, clicking pens or handling papers near the microphone.
- Agree on turn-taking, and say you will leave a short pause after each question.
- Ask for the correct spelling of their name, title and organization, and any terms they expect to use.
- Record ten seconds of them talking and listen back together if anything sounds off.
A researcher records a 50-minute interview with a hospital administrator. Before starting, she asks him to sit a hand's width from the microphone, explains that the recording will be transcribed, and notes the spellings of three colleagues and two software systems he mentions. She asks one question at a time and waits a beat after each answer before the next. When he cites figures, she repeats them back: "So that's eighteen months, one eight?" The transcript still needs a review pass, but the names are easy to fix from her list, there are no tangled overlaps, and the timestamps line up with clean pauses.
Mistakes and limits of good technique
- Good speaking habits cannot rescue a noisy room, a distant microphone or heavy echo; fix the setup too.
- Asking guests to speak unnaturally slowly makes them self-conscious and rarely helps much.
- Strong accents and dialects are not a delivery fault. Speak naturally; consistent distance and clear turns help more than trying to change an accent.
- Even careful delivery leaves some errors, especially in names and technical terms. Plan for a review pass.
- Switching languages mid-sentence is natural for many speakers but harder for most systems to transcribe accurately; if you want a clean single-language transcript, try to keep each answer in one language.
Speaking for mydubly transcripts and dubs
Several details of how mydubly processes a recording reward the habits above. The spoken language is detected automatically, so a recording mostly in one language is the straightforward case. The audio is split into chunks of about 30 seconds, with each cut placed at the quietest moment in the final few seconds of the window, so regular pauses give the cuts somewhere quiet to land instead of mid-word. A voice activity filter skips silence while keeping timestamps aligned to the original audio.
mydubly does not label speakers, so naming people as you hand over the conversation makes the transcript easier to follow. Subtitle files contain one single-line cue per recognized segment.
For dubbing, translated lines are fitted to the timing of the original speech: a line can start slightly early or end slightly late, and the synthesized voice may be sped up gently or slowed down. Translations are often longer than the source, so steady pacing with natural pauses gives the dubbed voice room to breathe, whereas rapid, unbroken speech leaves little room to fit. More on that output is on the AI dubbing page.
Next step
Before your next recording, read the guest-briefing list aloud to yourself and do a ten-second test. After the session, transcribe it with mydubly's audio to text tool, then check names against your list. A 50-minute interview costs 50 credits (5¢) to transcribe. For a complete interview workflow, see interview transcription.
Frequently asked questions
How fast should I talk for transcription?
At a natural conversational pace, or slightly slower if you tend to rush. Steady pacing with short pauses between thoughts matters more than a particular speed. Very slow, over-careful speech rarely improves accuracy and makes recordings tiring to listen to.
Does spelling names out loud help speech recognition?
The recognizer may transcribe the spelled letters rather than use them to correct the name, so do not expect the name to come out right automatically. Spelling still helps because the correct form is on the recording, and whoever proofreads the transcript can fix every instance quickly.
How do I stop people talking over each other in a recording?
Agree on turn-taking at the start, hand over by name, and leave a short pause after each answer. Ask participants to nod rather than say small acknowledgments. On remote calls, wait an extra beat before replying, because connection delay makes accidental overlaps more likely.
Should I remove filler words while speaking?
There is no need to force it; filler words are normal and often removed in a clean-verbatim transcript anyway. Replacing them with silent pauses when it comes naturally makes both the audio and the transcript easier to follow, but trying too hard can make speakers stilted.
What microphone distance should I keep?
For a typical desk or podcast microphone, around a hand's width is a common starting point. The exact number matters less than keeping it consistent: drifting closer and farther changes level and tone and adds room echo at the far end.
Will a strong accent hurt my transcript?
It can affect accuracy, depending on how well the recognizer's training data covers that accent. Speak naturally rather than trying to change your accent; clear distance, steady pacing and one speaker at a time do more to help. Review names and technical terms afterwards in any case.