What the comprehensible input idea claims
Comprehensible input is a term from the linguist Stephen Krashen, who set out his Monitor Model of second language acquisition in books and papers from the late 1970s into the 1980s. The model has five linked hypotheses. The one learners hear most about is the input hypothesis: people acquire a language by understanding messages in it, and they progress when the input is slightly beyond their current level, which Krashen wrote as i+1.
Two other parts of the model shape how the idea gets applied. The acquisition-learning distinction holds that consciously studied rules are a different system from the subconscious knowledge that produces fluent speech, and that only the second drives fluency. The affective filter hypothesis proposes that anxiety and boredom block input from being taken in, which is why advocates favor relaxed, interesting material over drills.
Put plainly, the claim is that the core activity of language learning is listening to and reading things you understand, thanks to context, pictures and the words you already know. Speaking ability, on this view, emerges as a result of that input rather than as its cause.
Where researchers push back
The input hypothesis has been influential in language teaching and heavily debated in second language acquisition research. The main criticisms are worth knowing, because each one changes how you should use video.
- Testability. Critics, Barry McLaughlin among the early ones, argued that the key terms are hard to define or measure. Nobody can say precisely what a learner's level is, so i+1 works as a metaphor rather than something you can test.
- Output. Merrill Swain studied students in Canadian French immersion programs who had received years of meaningful input yet still produced noticeably non-native French. She proposed the output hypothesis: having to speak and write pushes learners to notice gaps that comprehension lets them skip over.
- Interaction. Michael Long's interaction hypothesis emphasizes conversation in which people negotiate meaning, asking for clarification and rephrasing, as a central way input becomes understandable and useful.
- Noticing. Richard Schmidt's noticing hypothesis argues that learners need to consciously register features of the input, such as a verb ending or a particle, before those features are acquired.
- Explicit teaching. Many researchers dispute that consciously learned grammar can never feed into fluent use, and the role of instruction remains an active debate.
A fair summary is that large amounts of understandable input are broadly accepted as necessary, while the claim that input alone is sufficient is the contested part. For a learner, that suggests making comprehensible video the backbone of your study and adding speaking, writing and some deliberate attention to form alongside it.
How to judge whether a video is at your level
"Slightly beyond your level" needs a working definition before it helps you choose anything. Vocabulary coverage is the most practical one. Reading research associated with Batia Laufer and Paul Nation is often summarized as needing to know somewhere around 95 to 98 percent of the running words in a text for comfortable comprehension, and listening studies have proposed similar or somewhat lower thresholds. The figures vary with the test and the task, so treat them as a rough guide rather than a rule.
Turned around, those figures mean that a text where you meet more than a few unknown words in every hundred quickly stops being comprehensible without help. You can measure that on a transcript in five minutes, which is far more reliable than a feeling from the first minute of watching.
Vocabulary is not the only factor. These also move a video up or down in difficulty:
- Visual support. A cooking video shows what the words mean; two people talking at a desk show nothing.
- Speech rate and clarity. Scripted narration is usually slower and more carefully articulated than unscripted conversation.
- Topic familiarity. A subject you know well in your own language lets you predict what comes next.
- Number of speakers. One presenter is easier to follow than a group talking over each other.
- Accent and register. Regional accents, slang and fast casual speech are hard even when the vocabulary is simple.
- Videos made for learners in the target language
- Slow, clear and heavily visual. Often comprehensible from early levels.
- Tutorials, vlogs and cooking channels
- Natural speech with strong visual support. A common next step after learner material.
- Interviews and talks
- One or two speakers, more abstract vocabulary, little help from the picture.
- TV drama and films
- Fast dialogue, slang, several speakers and music under the speech.
- Panel shows, comedy and street interviews
- Overlapping voices and jokes that depend on culture and wordplay.
Making a hard video comprehensible with a transcript
A video slightly above your level can often be brought down to the right level with preparation. The transcript is what makes that possible, because it lets you deal with the unknown parts before and after watching instead of in real time.
- Pre-read. Read the target-language transcript of a section before you watch it. Listening then becomes recognizing words you have just seen rather than decoding unfamiliar sound.
- Gloss sparingly. Look up only the words that block the meaning of a sentence and write short glosses beside them. Leave the words you can guess.
- Use a translation once, for the gist. If a passage is still opaque, read the translated transcript for those lines, then return to the original. Don't watch with the translation open, or you will read instead of listen.
- Watch in sections. Use the timestamps to divide a long video into two- or three-minute parts and make each one comprehensible before moving on.
- Stay narrow. Krashen also proposed narrow listening: staying with one speaker, series or topic for a while. Vocabulary and voice keep recurring, so each new video is easier than the last.
Fading the support from subtitles to none, and the full week-long routine around a single video, are covered in learning a language with videos, so they are not repeated here.
Worked example: bringing a cooking channel down to your level
Suppose an intermediate learner of Spanish picks a 12-minute video from a Mexican cooking channel she enjoys. She makes a transcript for 12 credits (1.2¢), reads a stretch of about 100 words from the middle and underlines 11 she doesn't know. That is too many to follow comfortably, so instead of watching and hoping, she glosses the 20 words across the transcript that block meaning, pre-reads each section, and watches with Spanish subtitles. Two weeks later, the next video from the same cook has five unknown words in a similar sample, because the ingredients, the cooking verbs and the cook's habitual phrases keep coming back.
Three decisions made this work. She measured instead of guessing, so she knew the video needed scaffolding rather than abandoning. She glossed only blocking words, which kept the preparation to about twenty minutes. And she stayed with one channel, which is narrow listening in practice. Had the count come out at 25 unknown words per hundred, the right call would have been to shelve the video for a few months and come back to it.
Running a comprehensible input session, step by step
- Shortlist three or four videos on topics you care about, each under 15 minutes.
- Make a transcript and target-language subtitles for each.
- Sample each transcript: read about 100 running words from the middle and count the words you don't know, ignoring names.
- Sort the videos. As a rule of thumb, up to about four unknown words per hundred means watch as it is; roughly five to fifteen means scaffold it; much more than that means shelve it for now.
- For the scaffolded videos, gloss the blocking words and pre-read the first section.
- Watch section by section with target-language subtitles, then watch the same section again without them.
- Note any phrase worth keeping, with its timestamp, for your vocabulary list.
- Re-test a shelved video every month or so. Watching its count fall is a concrete measure of progress.
Mistakes that make input less useful
- Staying too comfortable. Material you understand completely builds fluency and is enjoyable, but it brings few new words. Mix in some stretch material.
- Treating noise as input. A far-too-hard video playing in the background exposes you to sound, not to understanding. Comprehension is the whole point of the word comprehensible.
- Reading the translation instead of listening. A translated transcript that is open the whole time turns a listening session into a reading session in your own language.
- Skipping output entirely. Given the criticisms above, add regular speaking or writing, even ten minutes with a tutor or language exchange partner.
- Trusting every transcript line. Recognition can mishear slang and names, and transcripts tend toward tidy text that drops fillers; see verbatim vs clean verbatim transcription for why the text and the audio sometimes differ.
- Counting only hours. An hours log is useful, as extensive listening explains, but coverage tells you whether those hours are at the right level.
Using mydubly to prepare input
mydubly produces the materials this approach relies on from a video or audio file on your device: a plain transcript, a timestamped transcript, and SRT and VTT subtitles in the spoken language, which is detected automatically. You can add a translated transcript in one of the other supported languages in the same job. Transcript mode costs 1 credit per minute with a 5-credit minimum per file, so the 12-minute video above costs 12 credits (1.2¢), and files can run up to 2 hours.
A few boundaries matter for learners. mydubly works on files, not links, so you need a copy of the video you are allowed to download. One job produces one subtitle track, so run the file again if you also want subtitles in a second language. Each subtitle cue is a single recognized speech segment on one line, which suits stepping through a video sentence by sentence. The 21 supported languages are listed on the subtitle generator page; Bengali, Tamil, Indonesian and Ukrainian, for example, are not among them. The SRT generator page explains the subtitle file itself.
Try it on a video you have been avoiding
Pick one video you set aside because it felt too hard. Make a transcript with the subtitle generator, run the 100-word sample, and decide whether it needs scaffolding or shelving. Once you have a few videos at the right level, building vocabulary lists from transcripts turns the words you glossed into ones you remember.
Frequently asked questions
Is comprehensible input the same as watching TV in the language I'm learning?
Only if you understand most of what you watch. A drama far above your level is input in the sense that sound reaches your ears, but much of it is not comprehensible, so little of it becomes acquired language. Watching TV becomes comprehensible input when the show suits your level, or when you prepare difficult episodes with a transcript first.
Can complete beginners learn from comprehensible input?
Yes, but not from ordinary native content. Beginners need material built for them: teachers or creators who speak slowly, use gestures, drawings and objects, and repeat key words many times. Transcripts help less at this stage because almost every word is unknown. Native videos with strong visual support become useful once you know a few hundred common words.
Do subtitles in my own language count as comprehensible input?
They make the video understandable, but the understanding comes mainly from reading your own language, so your listening gets little practice. A better compromise is to read the translated transcript for a hard passage once, then watch that passage with subtitles in the language you are learning, so the meaning you got from the translation attaches to the sounds.
How do I count unknown words in a Chinese or Japanese transcript?
Both languages are written without spaces between words, so a simple word count doesn't work. Use a reading tool or word segmentation tool that splits the text into words, or count unknown items across a fixed number of lines instead. The exact number matters less than comparing videos the same way each time.
Does comprehensible input mean I should never study grammar?
That is close to Krashen's own strong position, but many researchers and teachers disagree. A common middle path is to spend most of your time on understandable input and add short, focused grammar study when you notice a pattern you keep misunderstanding, such as an aspect distinction or a case ending. Use the grammar to make the next video clearer.