What an accent actually changes in the audio
An accent is a consistent pattern of pronunciation: which vowels a speaker uses, whether certain consonants are dropped or softened, where stress falls in a word, and how pitch rises and falls across a sentence. A speech model never sees the word "water"; it sees a spectrogram, a picture of energy across frequencies over time. When a speaker's vowels sit in a different place on that picture than the model expects, the acoustic evidence points toward a different word.
Most of the time the model recovers, because it also uses context. If the sound is ambiguous between "ship" and "sheep" but the sentence is about a ferry, the language-model side of the system pulls toward "ship". Errors appear when the acoustic shift is large and the context is weak, which is why short phrases, lists, names and numbers suffer most from accented speech.
- Vowel shifts: the same written vowel lands in a different acoustic region, so near-homophones get swapped.
- Consonant changes: dropped final consonants, a tapped or rolled r, or th pronounced as t or d.
- Stress and rhythm: syllable-timed speech (common in speakers of Spanish, French, Hindi or Turkish) can blur word boundaries for a model trained mostly on stress-timed English.
- Intonation: rising or falling pitch patterns affect where the model thinks one phrase ends and the next begins.
Why training data decides which accents work well
End-to-end models learn pronunciation implicitly from paired audio and text. If a training set contains thousands of hours of one accent and a handful of another, the model builds a detailed picture of the first and a rough sketch of the second. The original Whisper paper reports that English made up the largest share of its weakly supervised training audio, and that recognition quality in each language tracked how much training data that language had. The same logic applies within a language: accents common in podcasts, broadcast media and online video are learned well, and accents rarely heard in that material are learned less well.
This is a measured effect, not a hunch. A widely cited 2020 paper in PNAS by Koenecke and colleagues found substantially higher error rates for Black American speakers than for white American speakers across several commercial speech recognition services. The lesson generalizes: a model's accuracy on "English" is really an average across the accents it was trained on, and your speakers may sit well above or well below that average.
Non-native speech is a different problem from regional accents
Regional accents are stable systems shared by millions of speakers, and large models have often heard plenty of them. Non-native speech is less predictable. A speaker's first language shapes which sounds they substitute, and that pattern varies from person to person and even from minute to minute as they relax or tire. Non-native speakers also pause mid-phrase, restart sentences and choose less common word orders, which weakens the contextual guesses a model relies on.
Suppose a speaker whose first language is Spanish says "We need a cheap option for the pilot." Spanish has no separate short i vowel, so "cheap" and "chip" can sound alike, and the transcript reads "We need a chip option for the pilot." Every word is a real English word, so a spell checker will not catch it; only a reader who knows the topic will.
Regional vocabulary and dialect words
Pronunciation is only half the story. Dialects bring words and spellings the model may rarely have seen written down. Indian English "prepone", Scottish "loch" and "outwith", South African "robot" for a traffic light, or Irish "craic" are all standard in their regions. A model may hear them correctly and still write the more common neighbor, such as "lock" for "loch" or "crack" for "craic", because the common word is far more likely in its training text.
Place names, local institutions and family names are the hardest category. They are often absent from training text entirely, so the model spells them phonetically or swaps in a better-known name. That is why a list of names prepared before you transcribe saves more time than any other single step.
Where modern models handle accents well
It is worth being clear about the good news. Large multilingual models trained on varied web audio are far more tolerant of accents than older systems that relied on a fixed pronunciation dictionary for one "standard" accent.
- Widely spoken regional accents of major languages, such as Indian, Nigerian, Australian or Scottish English, or Mexican and Castilian Spanish, are usually transcribed well when the recording is clean.
- Long, connected speech gives the model enough context to correct individual sound errors on its own.
- Mild non-native accents in a clear recording often produce transcripts that need only light correction.
- Accent matters much less than recording quality; a strongly accented speaker on a close microphone usually beats a neutral accent in an echoing room. The guide on how to improve transcription accuracy covers the recording side.
Limits: what you cannot fix after recording
Some accent problems cannot be undone by any setting or tool. Being honest about them helps you budget proofreading time.
- You cannot retrain or fine-tune a hosted model on your speaker's voice, so a consistent mis-hearing will repeat throughout a file.
- Heavily dialectal speech that differs a lot from the written standard may be partly normalized toward standard spelling, partly misrecognized, or both.
- A strong accent combined with a short clip can confuse automatic language identification, so the model may transcribe in the wrong language. The article on spoken language identification explains why.
- When a speaker mixes a regional language with English, recognition of the switched-in words is often weak; see mixed-language videos.
- Errors compound downstream: a mis-heard word becomes a mistranslated subtitle and, in a dub, a wrong sentence spoken aloud.
Practical steps for accented recordings
You get the most benefit by aiming effort at the error types accents actually cause, rather than rereading the whole transcript line by line.
- Before transcribing, write down the names, places, products and dialect words you expect to hear, spelled correctly.
- Run the file and skim the transcript for those words first, using search; fix every variant spelling in one pass.
- Search for near-homophones typical of the accent, such as ship and sheep, live and leave, or three and tree.
- Listen back only where the text reads strangely, jumping straight there using the segment timestamps.
- Check numbers against the audio, since accent shifts on "thirteen" and "thirty" or "fifteen" and "fifty" are easy to miss.
- If one speaker is consistently hard to transcribe, consider re-recording key passages, especially for content that will be translated or dubbed.
For a deeper review workflow, the guide to proofreading an AI transcript shows how to prioritize what to check.
How mydubly handles accented speech
mydubly's transcription runs on OpenAI's open-source Whisper model, with Whisper large-v3-turbo as the default deployment, so its behavior on accents follows Whisper's strengths and weaknesses described above. You do not choose the spoken language or an accent; mydubly detects the spoken language from the audio, and you only pick the output language: the spoken language for an untranslated transcript, or another language for a translation.
The output gives you what an accent-focused review needs: a timestamped transcript and SRT or VTT subtitles, so you can jump to any suspicious line. Transcripts cost 1 credit per minute, so a 45-minute interview with a strongly accented guest costs 45 credits, or 4.5 cents, which makes it cheap to transcribe first and decide how much correction a file needs. Language pages such as English video to text and Spanish video to text list what each language supports.
Where to go next
If your recordings feature speakers with strong regional or non-native accents, start with a short test: upload a few minutes to video to text, compare the transcript to what you hear, and note which words the accent breaks. That small sample tells you how much proofreading the full file will need before you commit to subtitles, a translation or a dub.
Frequently asked questions
Does Whisper understand every English accent equally well?
No. Whisper is trained on a large and varied set of web audio, so it handles many accents well, but accents that were rare in that data still produce more errors. Widely heard accents from broadcast and online media tend to do best.
Should I ask speakers to change their accent for transcription?
No, and it rarely works anyway. Ask for things that help every speaker instead: a close microphone, one person talking at a time and a moderate pace. Then correct the predictable accent errors in a targeted proofread.
Why does the transcript spell local place names wrong even when they sound clear?
Rare names barely appear in training text, so the model either spells them phonetically or substitutes a more common name that sounds similar. A pre-made list of names lets you fix every variant with search and replace.
Can a strong accent make the model pick the wrong language?
It can, especially when the speech sample is short or the speaker mixes languages. A long stretch of clear speech gives the language detector more evidence; if the transcript comes back in the wrong language, that is the first thing to suspect.
Is accented speech more expensive to transcribe in mydubly?
No. Transcription is priced by duration only, at 1 credit per minute with a 5-credit minimum per file, regardless of who is speaking or how they sound.