Three methods compared
- Typing by hand
- Accurate if you're careful, but typically several hours of work per hour of video.
- Platform captions
- Free if the video is on YouTube or a meeting tool with transcripts on — quality varies, and you need access to that platform's copy.
- Automatic transcription
- Minutes per file, timestamps included, and you can edit the result. Costs a little per minute.
Method 1: automatic transcription (fastest)
- Open the video to text converter and switch Full translation output off.
- Drop in the video. Your browser extracts the audio; the video itself isn't uploaded.
- Select the spoken language for an untranslated transcript, or another language for a translated transcript too.
- Download the plain transcript, the timestamped transcript, or SRT subtitles.
A 45-minute recording costs 45 credits (4.5¢); a 2-hour one 120 credits (12¢).
Method 2: use existing platform captions
On YouTube, open the video, click … under the player and choose Show transcript. Zoom, Teams and Meet can produce transcripts if they were enabled before the recording. These are handy when available, but you can't always download them, and quality depends on the platform's settings.
Method 3: type it yourself
Manual transcription still makes sense for very short clips or legal-grade verbatim work. Slow the playback to 0.75×, use a foot pedal or keyboard shortcuts, and type in short bursts. Even then, starting from an automatic draft and correcting it is usually faster.
Getting an accurate transcript
Accuracy depends far more on audio than on the tool. Background music, room echo and people talking over each other cause most errors. See how to improve transcription accuracy for recording tips.
Which download fits the job
A transcript-mode run gives you four files from the same transcript. They differ in what you can do with them:
- Plain transcript
- Reading, quoting, or pasting into a document or notes app. It holds the translated text if you picked another language, or the untranslated text if you picked the spoken language.
- Timestamped transcript (.txt)
- Finding moments, citing "at 14:32", or giving an editor a paper edit to cut from.
- SRT
- Captions on YouTube, Vimeo or LinkedIn, and importing into Premiere Pro or DaVinci Resolve.
- VTT
- Captions in an HTML5 player on your own site; the VTT generator page shows the file structure.
Download every file you might need before you close the result screen, because outputs are deleted within 30 minutes of the job ending. With a target language picked, Download transcript gives the translation and Download original transcript gives the wording as spoken, so one run covers both texts.
From recording to usable text
- Check the audio before uploading. Scrub to three random points and listen. If speech is buried under music, or the file carries two audio tracks (screen recorders sometimes capture the microphone and system sound separately), export a version with just the dialogue. See multiple audio tracks in video files.
- While the job runs, jot down rough times when each person starts speaking. Transcripts don't label speakers, and those notes make the attribution pass much quicker.
- Run transcript mode. Select the spoken language itself for untranslated text in that language.
- Download the timestamped transcript and the SRT.
- Proofread against the video, starting with names, numbers and anything you plan to quote. The method in how to proofread an AI transcript keeps this fast.
- Add speaker names at the turns you noted, then save a clean copy for sharing.
How much to edit depends on where the text ends up. A transcript published beside the video for accessibility or search should stay close to verbatim, with only the false starts removed that make it hard to read. Quotes in an article can be lightly tidied for readability as long as the meaning doesn't change, and the untouched timestamped file stays your record of what was actually said.
Troubleshooting a disappointing transcript
- Whole sentences missing: usually overlapping speech, music louder than the voice, or a speaker far from the microphone. Listen at those timestamps; transcription missing words covers each cause.
- Text switching language mid-recording: the spoken language is detected automatically, and a video that changes language partway can be handled inconsistently. Split the file where the language changes and transcribe each part.
- Every brand name or acronym is wrong: recognition guesses unfamiliar words from their sound. Keep a list and fix them with search and replace in both the transcript and the SRT so they stay consistent.
- Captions drift away from the picture in your editor: you probably transcribed a different cut. Any trim made after transcription shifts every later cue, so transcribe the exact export you publish.
- Song lyrics appear in the text: sung vocals in background music can be transcribed like speech. Delete those lines, or cut long music sections before uploading.
- The transcript is nearly empty: the audio track may be silent or muted. Play the file in a local player and check the volume meter before rerunning.
Worked example: a 50-minute recorded panel
A conference organiser has a 50-minute panel with a moderator and three guests, recorded as an MP4 from the venue's camera. Transcript mode costs 50 credits (5¢). They download the timestamped transcript and the SRT. Because speakers aren't labelled, they mark each turn using the timestamps and the video. They pull three quotes for a recap blog post and verify every quote against the recording at its timestamp before publishing. The SRT goes onto the YouTube upload as captions. Partners abroad want a Spanish version, so a second run with Spanish as the target adds another 50 credits (5¢) and produces Spanish subtitles too. If your files come from phones or cameras, the MP4 to text page lists format quirks worth checking, and webinar translation covers the same job for recorded online events.
Frequently asked questions
Can I transcribe a video without uploading it?
With mydubly, only the audio leaves your device — the video file itself stays in your browser.
Does the transcript include speaker names?
No. Transcripts are timestamped but don't label speakers.
Is a transcript the same as a subtitle file?
No. A transcript is continuous text meant for reading; a subtitle file splits the same words into short timed cues meant to appear over the video. Transcript mode gives you both from one run.
How do I find one moment in a long video?
Search the timestamped transcript for a word you remember from that moment, read the time next to it, and jump straight there in your player.
What if my video is longer than two hours?
Split it into parts under two hours and transcribe each. Timestamps start from zero in every part, so add each part's start time when you cite a moment from the full recording.