What dictation trains that ordinary listening doesn't
When you listen for meaning, your brain fills gaps with guesses, and you can understand a sentence while missing half its words. Dictation removes that safety net. Writing every word forces you to resolve word boundaries, catch unstressed words such as articles and auxiliaries, hear inflectional endings, and connect sounds to spellings. Teachers often call it a bottom-up exercise, because it works on decoding the sound stream rather than on grasping the message.
It is also diagnostic. A gist question tells you whether you understood; a dictation shows exactly where hearing broke down. A learner of English who writes "I want go" for "I wanted to go" has a gap with weak forms and past-tense endings that no comprehension score would reveal.
Four ways to run a dictation
- Full dictation
- You write every word of a short passage. The most demanding version and the most informative.
- Partial dictation
- You get the text with gaps and fill in the missing words while listening. Easier, and good for targeting one feature such as verb endings.
- Dictogloss
- Associated with Ruth Wajnryb's work on grammar dictation: learners hear a short text at normal speed, note key words, then reconstruct it, often in groups. It trains grammar and meaning more than exact transcription.
- Running dictation
- A classroom activity: a text is posted away from the desks, one learner reads and memorizes a line, then returns and dictates it to a partner. It mixes reading, speaking, listening and writing.
For self-study, full and partial dictation are the practical choices, and a transcript is what makes them possible without a teacher.
Choosing audio for dictation
- Length: 30 to 90 seconds for a full dictation. Longer passages turn into a test of stamina.
- One clear speaker with little background noise or music under the voice.
- Level: material whose general meaning you follow but whose details you can't yet catch reliably. If you can't follow the gist, you are guessing, not decoding.
- Scripted first, natural later. Narration is a gentle start; unscripted speech has the reductions and hesitations real listening involves.
- Content you don't mind hearing a dozen times, because you will.
- A source you can get as a file, so you can loop segments and make a transcript.
Taking several passages from one longer recording works well. You get variety in content while the speaker's voice and habits stay constant.
A self-study dictation session
- Listen to the whole passage once without writing, just for the gist.
- Play it one sentence or clause at a time, pausing to write. A player with an A-B loop function, or the transcript timestamps, makes this easy.
- Replay each segment up to three or four times. When you can't get a word, leave a gap rather than guessing wildly, and mark words you are unsure of.
- Listen to the whole passage again and fix anything that now sounds different.
- Only now open the transcript and compare line by line.
- Mark every difference and classify it using the categories below.
- Listen once more while reading the transcript, paying attention to the places you missed.
- Repeat the same passage a few days later and compare your scores.
Typing is fine for most languages and makes comparison easier: put the transcript in one spreadsheet column and your version in the next, one segment per row.
Checking your version against an AI transcript
The transcript is your answer key, which raises an awkward fact: an AI transcript is a good answer key but not a perfect one. Before counting a difference as your mistake, decide which version is right.
- If you wrote something different, replay the segment. If the audio clearly supports your version, record it as a transcript error, not yours.
- Machine transcripts tend toward a clean style in which fillers, repeated words and false starts are often dropped. If you wrote a hesitation the transcript lacks, don't count it as an addition; verbatim vs clean verbatim transcription explains the difference.
- Punctuation and capital letters are added by the model. Ignore them unless you are practicing them.
- Numbers may appear as digits or as words. Accept either.
- Names and rare terms are where recognition errs most often, so check them against a reliable source.
If you plan to reuse a passage many times, or give it to students, proofread the transcript against the audio once before using it as a key. Ten minutes of checking saves you from memorizing the model's mistakes.
Scoring mistakes by type
A count of errors gives you a number to track. Classifying them tells you what to practice next.
- Missed word
- A word in the audio you left out. Often an unstressed article, auxiliary or preposition.
- Wrong word
- You wrote a different word, often one that sounds similar or a more familiar word that seemed to fit.
- Word boundary
- You split or merged words wrongly, hearing one word where there were two or the reverse.
- Ending or form
- Right word, wrong form: a missing plural, tense or agreement ending.
- Spelling only
- You heard the right word but spelled it wrong. A writing issue rather than a listening one.
- Added word
- You wrote a word that was not said, often to make the sentence feel grammatical.
For a single score, count every error except spelling, divide by the number of words in the passage and multiply by 100, which gives errors per hundred words. This is the same idea as word error rate, the standard measure for speech recognition, described in word error rate explained. Track spelling separately if it matters to you.
Suppose an intermediate learner of French takes a 60-second clip from a news summary, about 150 words long. Her first attempt has 13 errors besides spelling: four missed small words such as de, en and y, five wrong endings, two boundary errors and two wrong words. That is roughly 9 errors per hundred words. A week later, on the same clip, she makes 6, about 4 per hundred, and a new clip from the same presenter comes in at 7 per hundred.
The classification changed what she practiced. Many French verb and plural endings sound identical, so her five ending errors were mostly a grammar problem, not a hearing problem, and she reviewed agreement rules. The missed small words were a genuine listening gap, so she made partial dictations that blanked out every short function word.
Progressing over the weeks
- Stay with a passage until you score a few errors per hundred words, then move on.
- Lengthen passages gradually, from 30 seconds toward two minutes.
- Move from scripted narration to unscripted speech.
- Vary speakers: different ages, speeds and regional accents.
- Allow fewer replays, from four per segment down to two, and dictate longer chunks at a time.
- Target weaknesses with partial dictation, for example a gapped text that removes every past-tense verb.
- Keep a simple log with the date, passage, length, errors per hundred words and the most common error type.
Limits of dictation as an exercise
- It is slow. A one-minute passage can take half an hour to dictate and check, so it complements large amounts of easier listening rather than replacing it; see extensive listening.
- It trains form more than meaning. You can score well on a passage and still miss an implication or a joke.
- In languages with irregular spelling, such as English and French, dictation also tests spelling, so keep that category separate.
- In Chinese and Japanese, writing characters by hand is a separate skill from hearing. Many learners type with pinyin or romaji input and choose characters, which tests recognition instead.
- Transcript errors can mislead you, especially with strong accents or music under the speech.
- Material that is too hard turns into frustration and guesswork.
Making dictation keys with mydubly
mydubly's transcript mode produces the answer key. You add an audio file such as MP3, M4A or WAV, or a video file, from your device and get a plain transcript, a timestamped transcript, and SRT and VTT subtitle files in the spoken language, which is detected automatically. Each subtitle cue is one recognized speech segment, which often makes a convenient dictation unit. You can add a translated transcript to check meaning after you finish, which is useful but separate from the dictation itself.
Pricing favors transcribing a whole recording once rather than many tiny clips. Transcript mode costs 1 credit per minute with a 5-credit minimum per file, so a single 60-second clip still costs 5 credits (0.5¢), while a 30-minute episode costs 30 credits (3¢) and yields dozens of passages. mydubly doesn't compare your dictation with the transcript or score it; the comparison is yours to do. The audio to text page lists supported formats, and the timestamped transcript format page shows what the key looks like.
Start this week
Choose a recording with one clear speaker, transcribe it with audio to text, and pick a 45-second passage from the middle. Dictate it with the eight-step method, classify every mistake, and note your errors per hundred words. Repeat the same passage on Friday. If you want a quick way to check a transcript before trusting it as a key, read how to proofread an AI transcript.
Frequently asked questions
Should I type or handwrite a dictation?
Either works for listening practice. Typing makes comparison against the transcript quicker, because you can align both versions in a spreadsheet. Handwriting adds spelling and handwriting practice, which matters in some exams and in languages with characters. Pick the mode that matches what you need to do in real life, and keep it consistent so scores stay comparable.
How many times should I replay each segment?
Three or four replays per segment is a common ceiling. After that, extra replays rarely reveal a word you couldn't catch, and you start to imagine sounds that fit your guess. Leave a gap, finish the passage and let the transcript show you what was there. Reduce the number of replays as your accuracy improves.
Is dictation useful for beginners?
Yes, in small doses. Beginners do better with partial dictation, where most of the text is given and only a few words are missing, or with very short passages of five to ten seconds. Full dictation of natural speech is usually too demanding at first, because so many words are unknown that the exercise turns into guessing.
Can I practice dictation with song lyrics?
You can, but songs make poor dictation material. Singing stretches vowels, changes stress and hides words under instruments, and lyrics often bend grammar to fit the melody. Speech recognition also struggles with singing, so an AI transcript is a weaker key; see translating song lyrics in videos for why. Use speech for scored practice and songs for fun.
How do I turn a transcript into a partial dictation?
Copy the plain transcript into a document and replace the words you want to practice with blanks: every seventh word for general practice, or a whole category such as articles, prepositions or past-tense verbs for targeted work. Keep the original as the answer key, and use the timestamps to find each sentence when you replay it.