Why mine words from a transcript instead of while watching
Pausing a video every time you hear an unfamiliar word breaks the listening and produces a list biased toward whatever you happened to catch. A transcript lets you separate the two jobs: watch for meaning first, then go through the text calmly and choose words on purpose. It also gives you three things a list made by ear can't: the exact spelling, the full sentence the word appeared in, and a timestamp that leads back to the audio.
That last point matters more than it seems. A word you learn together with its sound and its sentence is easier to recognize in fast speech later than a word learned as a spelling and a translation. A timestamped transcript keeps all three together in one file.
Finding what is frequent in one video
Two kinds of frequency are useful. General frequency is how common a word is in the language as a whole, and frequency lists built from large collections of text and film subtitles are available for many languages. Local frequency is how often a word appears in the video you are studying. A word that is rare in general but appears nine times in a video about car repair is worth learning if you plan to watch more videos about car repair.
You don't need special software to count local frequency:
- Paste the plain transcript into a text editor and replace punctuation with spaces.
- Replace spaces with line breaks so there is one word per line.
- Paste that column into a spreadsheet, convert it to lowercase, and count each form with a pivot table or a COUNTIF formula.
- Sort by count and skim from the top.
The top of the list will be articles, pronouns and prepositions; skip past them. Two complications are worth knowing. Inflected languages spread one word across many forms, so Spanish hablo, hablas and hablaron count separately until you group them under the dictionary form, a step called lemmatization that some free text-analysis tools automate. Chinese and Japanese are written without spaces between words, so splitting on spaces doesn't work, and you need a word segmentation tool, or a reading tool that segments for you, before counting.
Choosing which words earn a place
A long video can contain hundreds of words you don't know, and trying to learn all of them is the fastest way to abandon the list. A short filter helps:
- It appears more than once in the video, or you have met it elsewhere recently.
- It sits in the general-frequency range you are working on, not far beyond it.
- You would use it, or expect to hear it again, given what you watch and talk about.
- You couldn't guess it confidently from context. Words you can guess are already being learned.
- It is not a name, a brand or a one-off technical term.
Collect chunks as well as single words. Collocations such as make a decision, fixed phrases, and verbs together with their usual preposition are often more useful than isolated words, and the transcript shows exactly how each was used.
A realistic target is 10 to 20 items per video, chosen from a few hundred candidates. The rest are not lost: the common ones will come back in your next video.
Sentence mining: keep the whole line
Sentence mining is the habit, popular among self-taught learners, of collecting whole sentences rather than words. The sentence goes on the front of a card with the target word highlighted or blanked out. A widely repeated rule of thumb is to pick sentences with only one unknown item, so the rest of the sentence supports the meaning, which echoes the comprehensible input idea of material just beyond your level.
Transcript segments make natural candidates, because each one is a stretch of speech a few seconds long. Copy the segment, note its timestamp, and if the line runs on, trim it to the clause that holds your word.
Suppose a learner of German studies a 20-minute video in which a mechanic walks through a car inspection. The transcript costs 20 credits (2¢). Her frequency count puts Reifen (tire), Bremse (brake) and prüfen (to check) near the top, each repeated several times. She first skips Hauptuntersuchung, the formal vehicle inspection, as too specialized, then keeps it after counting seven uses. She ends with 18 sentence cards, each holding the German line, an English gloss of the target word only, a four-second audio clip and the timestamp, such as 06:41.
Notice what she left out: dozens of words that appeared once, the mechanic's name, and several dialect expressions she couldn't confirm in a dictionary. Eighteen cards is about one evening's work, and every card links back to a moment she can replay.
Building flashcards with audio and timestamps
A useful card has four fields: the target-language sentence with the word marked, a short meaning in your language or a simple definition, an audio clip of the original line, and a source field with the video name and timestamp. The source field is what lets you check a doubtful card against the original weeks later.
Cutting the audio clips is the fiddly part. Two practical ways:
- In a graphical audio editor such as Audacity, open the audio, select from the segment's start time to its end time, and export the selection as an MP3.
- With ffmpeg on the command line, one command per clip does it: ffmpeg -i lesson.mp4 -ss 00:06:41 -to 00:06:45 -vn clip-0641.mp3 saves the four seconds starting at 6:41 as an MP3.
Add about half a second on either side of the segment times, because a clip cut exactly at the boundaries can lose the first or last sound.
Anki, a free and widely used spaced repetition app, can import cards from a CSV or tab-separated text file with one column per field. Audio goes into a field written as [sound:clip-0641.mp3], with the file itself placed in Anki's media folder. Anki also has a cloze note type, which hides a marked part of a sentence and suits sentence mining well. Import details change between versions, so check the current Anki manual before building a large file.
Spaced repetition without drowning in reviews
Spaced repetition software brings each card back just as you are likely to forget it, with the gaps growing each time you answer correctly. It builds on the spacing effect: since Hermann Ebbinghaus's memory experiments in the 1880s, psychologists have repeatedly observed that study spread out over time is retained better than the same amount crammed together. Paper methods such as the Leitner box system came before apps; Anki's long-standing scheduler derives from the SM-2 algorithm published for SuperMemo, and recent versions also offer an alternative scheduler called FSRS.
The practical risk is volume, not the algorithm. Every new card creates future reviews, so a learner who adds dozens of cards a day soon faces a review pile that takes longer every morning. Cap new cards per day at a number you can sustain on a bad day, rewrite cards you keep failing with a clearer sentence, and delete cards for words you no longer care about. A small deck you review daily does more than a large deck you abandon.
From transcript to deck in one sitting
- Watch the video once for meaning, without stopping.
- Get the timestamped transcript and the plain text.
- Run the local frequency count and skim the top of the list.
- Read the transcript and mark candidate words, using the filter above.
- Cut the list to 10 to 20 items.
- For each item, copy the shortest transcript segment that contains it and note the timestamp.
- Play every chosen line before making its card, and fix any recognition errors.
- Cut the audio clips, build the import file and import it into your flashcard app.
- Watch the video again a few days later and notice which words you now hear without effort.
Mistakes that make vocabulary lists stop working
- Copying a misrecognized word. AI transcripts can mishear rare words and names, and a wrong word on a card gets memorized wrongly. Step 7 exists for this; proofreading an AI transcript shows a quick way to check.
- Trusting a single-word machine translation. Out of context, a word often gets the wrong sense, like the river bank and the savings bank in English. Gloss from the sentence and a dictionary, not from translating the word on its own.
- Over-collecting. Hundreds of cards from one video creates a backlog, not vocabulary.
- Leaving out the sound. Cards without audio train you to recognize spellings, which is not the skill listening needs.
- Never rewatching. The deck and the video reinforce each other, and meeting the word again in context is part of learning it.
- Filtering out every small word. Frequency filtering skips particles and prepositions, but those whose use differs from your language deserve sentence cards of their own.
How mydubly fits a vocabulary workflow
mydubly's part in this workflow is producing the text. You add a video or audio file from your device and get a plain transcript, a timestamped transcript, and SRT and VTT subtitles in the spoken language, which is detected automatically, plus an optional translated transcript. Transcript mode costs 1 credit per minute with a 5-credit minimum per file, so the 20-minute video in the example costs 20 credits (2¢). Files can be up to 2 hours long, and the 21 supported languages are listed on the video to text page; Spanish video to text shows one language in detail.
mydubly doesn't count frequencies, choose words, cut clips or make flashcards. Those steps happen in a spreadsheet, an audio editor and your flashcard app. Transcripts are clean rather than strictly verbatim, so fillers and false starts may be missing, which rarely matters for vocabulary but explains small differences between the text and what you hear.
If you are just starting with a language, learn a language with videos describes a weekly routine that a vocabulary list slots into.
Start with one video you like
Choose a video under 20 minutes that you will happily watch three times, get its transcript from video to text and make no more than 15 cards. After a week, rewatch it without subtitles. If you study mostly from audio, the same method works for episodes; see learning a language with podcasts.
Frequently asked questions
How many new words should I take from one video?
For most learners, 10 to 20 items per video is sustainable, especially if you study a new video every few days. The limit is your daily review time, not the video. If your flashcard reviews regularly take longer than you planned, add fewer cards per video rather than skipping reviews, because missed reviews pile up quickly.
Should the back of a card show a translation or a definition in the target language?
Early on, a short translation is quicker and less likely to mislead. As your level rises, simple target-language definitions keep you thinking in the language and avoid false one-to-one equivalents. Many learners use both: a brief gloss in their own language plus the original sentence, which already shows how the word behaves.
Can I build vocabulary lists from Chinese or Japanese videos?
Yes, with an extra step. The transcript arrives in ordinary characters without word spaces, and Chinese output uses Simplified characters. Run the text through a segmentation or reading tool to split words and add pinyin or furigana if you want readings on your cards, then follow the same choosing and sentence mining steps.
Is it worth keeping a word I only heard once?
Usually not, unless you know it is common in general or important for something you do. A word heard once in one video is a weak candidate compared with one that recurs. If it is common, you will meet it again soon, and the second encounter is a better moment to make the card.
Can I share my flashcard deck with classmates?
Cards with your own example sentences are yours to share. Cards containing audio clips and transcript lines from someone else's video are copies of their work, so keep those decks for personal study unless the creator allows reuse. This is general information rather than legal advice; rules vary by country and by license.