Decide what you need before you start
Speed comes from knowing what to ignore, and you can only ignore things deliberately if you know what you are looking for. Before opening the file, write one or two sentences: what question should this recording answer for you? "Did the supplier commit to a delivery date?" leads to a very different review from "what are the main concerns staff raised?"
That question sets the depth of review:
- Triage
- Is this recording relevant at all? Skim the transcript only.
- Targeted
- Find and confirm specific content. Skim, flag, listen to flagged parts.
- Thorough
- Understand everything said. Skim, then listen to most of it at speed, with a full log.
Most reviews are targeted. Thorough reviews are for material you will rely on heavily, such as an interview at the center of a story or a hearing you must report accurately.
Why reading first is faster than listening
Listening is linear: a 2-hour recording takes 2 hours at normal speed, and you cannot see what is coming. Reading is not. Most people read considerably faster than people speak, and on a page you can jump, scan headings in your mind, and spot a name halfway down without processing everything above it.
A transcript gives you a map of the recording. You can see where the long monologues are, where the conversation turns, and where someone mentions the topic you care about. Even a raw machine transcript with errors is good enough for this; you are reading for structure and topics, not for exact wording. Correcting the transcript is a separate task, covered in proofreading an AI transcript, and it is usually unnecessary for review.
Skimming the transcript
Skimming is a technique, not just reading faster. Work through the timestamped transcript in one pass with your purpose written above it:
- Search first. Run a few searches for the names, terms and numbers your question depends on, and mark every hit. Recognition errors affect names most, so search for distinctive fragments as well as whole words.
- Read the first line of each passage and keep going unless it touches your question.
- Watch for signposts that mark turns in a conversation: "moving on", "next item", "the real problem is", "can I ask about", "to be honest". They often precede the substance.
- Notice questions. In interviews and hearings, the question tells you what the next few minutes are about.
- Look at the time labels. A long gap between two labels with little text can mean silence, music, or speech the recognizer missed. Mark these for a quick listen.
Do not stop to fix errors or take detailed notes during the skim. Your only job is to flag.
Flagging: a small set of marks
Use a handful of marks you can apply instantly, either in a copy of the transcript or in a separate log. For example:
- KEY for content that directly answers your question
- QUOTE for wording you may need exactly
- CHECK for anything that looks wrong, unclear or garbled in the transcript
- GAP for stretches where the transcript seems thin or missing
- LATER for interesting material outside today's purpose
Keep the scheme stable across recordings so a colleague can read your flags, and so you can search for them later. Flag generously during the skim; it is cheaper to listen to an extra minute than to miss something.
Listening to flagged parts at speed
Now open the recording and go only to the flagged moments. Start 15 to 30 seconds before each time label, because the context that explains a statement usually comes just before it.
Most media players and podcast apps let you raise playback speed while keeping the voice at its natural pitch. VLC, for example, has speed controls in its Playback menu, and many browser-based players offer a speed setting. Practical guidelines:
- Raise speed gradually. Many people find 1.25x to 1.5x comfortable for clear speech, and some go faster with practice.
- Slow down for dense content, strong accents, poor audio, overlapping speakers and anything flagged QUOTE.
- Listen at normal speed whenever tone matters, such as hesitation, sarcasm or an emotional answer. The words alone can mislead.
- Use headphones. Speed amplifies the effect of noise and room echo.
As you listen, resolve each flag: confirm it, correct the transcript wording in your log, or dismiss it.
Keeping a timestamp log
A log turns a review into something you, or someone else, can use later without repeating it. One line per item is enough:
- 00:14:20
- KEY: supplier says the new line is "about six weeks" away; hedged, no firm date
- 00:31:05
- QUOTE: exact words on safety checks, confirmed against audio
- 00:47:50
- CHECK: transcript says "fifteen units"; audio says "fifty units"
- 01:12:30
- GAP: two minutes of crosstalk, mostly inaudible; nothing relevant
Write the time, the mark and a note in your own words. Note what you confirmed in the audio, not just what the transcript says. A log like this also makes summaries easier to write and to check, because every statement already has a timestamp; how to build a summary from it is in summarizing a transcript.
A hypothetical reporter has a 110-minute interview with a former city official and needs every passage about one housing contract. She reads the transcript in about 15 minutes, searching for the contractor's name, "contract" and "tender", and flags nine passages plus two thin stretches. She listens to all eleven at 1.5x, dropping to normal speed for two answers where tone matters, which takes about 18 minutes. Her log has fourteen lines, including one transcript error on a figure. A final spot check of four random 2-minute windows finds nothing new. Total time is roughly 40 minutes, and every claim she uses is tied to a confirmed timestamp.
Sampling to check nothing was missed
Skimming relies on the transcript, and transcripts miss things. Finish with a short check that tests the parts you did not listen to:
- Pick several unflagged windows spread across the recording, for example one 2-minute window in each half hour, and listen at speed while reading along.
- Listen to the first and last few minutes. Important remarks are often made while people settle in or after someone thinks the formal part is over.
- Play every GAP flag, since missing text is where unexpected content hides.
- If a sample turns up something relevant that the skim missed, widen the review around it and consider listening to that whole section.
Sampling cannot prove a negative, but it tells you whether the transcript is a trustworthy map of this particular recording. If samples keep revealing content the transcript dropped, the audio is probably difficult, and a fuller listen is the safer choice. Common reasons transcripts skip speech are covered in why transcriptions miss words.
Mistakes and limits of transcript-led review
- Reviewing without a question. Without one, everything looks relevant and nothing gets skipped.
- Trusting the transcript for exact words, figures or names. Use it to find, use the audio to confirm.
- Listening too fast through the parts that matter. Speed is for getting to the content, not for absorbing it.
- Losing track of speakers. If the transcript has no speaker labels, note who is talking in your log as you listen.
- Skipping the sample check because the skim went well. A clean-looking transcript can still have a missing stretch.
- Reading tone from text. Irony, reluctance and emphasis rarely survive transcription, so listen to any passage where meaning depends on delivery.
Using mydubly for the transcript step
mydubly makes the transcript you skim. You choose a recording from your device, audio (MP3, WAV, M4A, AAC, OGG, FLAC) or video (MP4, MOV, WebM, MKV, M4V), up to 2 hours per file. The audio is cut into chunks of roughly 30 seconds at quiet moments and recognized with Whisper, with silence skipped while timestamps still refer to the original recording, so the [m:ss] labels in the timestamped transcript match the times in your player. You also get a plain transcript.txt and SRT and VTT files. See audio to text for audio and video to text for recordings with a picture.
mydubly does not flag, review, summarize or store recordings, and it does not label speakers. Results are deleted within 30 minutes of a job finishing, so download them first. For recordings longer than 2 hours, split the file and review the parts in order; the effects of chunking on very long audio are explained in transcribing long audio files. Keep the browser tab open while the job runs. Transcription is 1 credit per minute with a 5-credit minimum per file, so a 110-minute interview costs 110 credits (11¢).
Next step: try it on one recording you have been putting off
Take the long recording you have been avoiding, write your question at the top of a blank log, and transcribe it. Skim with the five marks, listen to the flags at a comfortable speed, and finish with the sample check. For interviews specifically, the interview transcription use case covers getting from recording to text.
Frequently asked questions
How fast can I listen without missing things?
It depends on the speaker, the audio quality and your familiarity with the topic. Many people are comfortable at 1.25x to 1.5x for clear speech and slow down for accents, crosstalk and important passages. Increase speed gradually and drop back whenever you find yourself replaying sections.
Is skimming the transcript enough on its own?
For triage, often yes. For anything you will quote, report or act on, confirm it in the audio, because transcripts can misrecognize names and numbers, drop quiet speech and lose tone. Treat the transcript as a map and the recording as the evidence.
How do I review a recording with several people talking?
Machine transcripts often have no speaker labels, so identify speakers as you listen to flagged sections and write their names or roles in your log. Overlapping speech is where transcripts are weakest, so flag crosstalk for listening rather than relying on the text.
Can several people share the review of one long recording?
Yes. Split the recording by time range, agree on the same flag marks and log format, and have each reviewer fill in their range. Ask one person to do the sample check across the whole recording at the end so the ranges are checked consistently.
What should I do if the transcript looks badly wrong?
If whole passages look garbled, the audio is probably difficult: noise, distance from the microphone, music or overlapping voices. Use the transcript only for rough navigation, listen to more of the recording, and consider whether the source audio can be improved before transcribing again.