What people turn videos into text for
- Meetings
- Minutes and action items from Zoom, Teams and Meet recordings — see meeting transcription.
- Lectures
- Searchable notes from recorded classes — see lecture transcription.
- Content
- Blog posts, show notes and quotes from videos you've published.
- Accessibility
- Text and captions for viewers who can't hear the audio.
Priced for long recordings
Transcription costs 1 credit per minute — 60 credits (6¢) for an hour, 120 credits (12¢) for two. Adding a translated transcript costs nothing extra.
How it works
- Drop in a video. Your browser reads it and pulls out only the audio — the video file never leaves your device.
- The audio is processed in 30-second chunks: speech is transcribed and the spoken language is detected automatically.
- Download the results. Audio and text on our servers are deleted after delivery, within 30 minutes at most.
What you get
- Transcript
- Plain text (.txt) of what was said
- Timestamped transcript
- Every line prefixed with its time, e.g. [12:04]
- Subtitles
- SRT and VTT, in the language you select (select the spoken language for untranslated subtitles)
Supported files
- Video: MP4, WebM, MOV, MKV, M4V — whatever your browser can read.
- Length: up to 2 hours per file.
- Languages: speech in 21 languages is recognised automatically; translation and voices cover the same 21.
Privacy
Your video is never uploaded — only its audio is, over HTTPS in small chunks. Audio, transcripts and voice files are deleted after delivery, within 30 minutes at most, and are never used to train models. Your account keeps only the file name, length, languages and credits used. Details in the privacy policy.
Good to know
- No speaker labels — transcripts don't say who is speaking.
- Files only: no links from YouTube or social apps, and no live audio.
- Subtitles come as SRT/VTT files rather than burned into the picture.
Picking the right download for the job
Transcript mode gives four files, and each suits a different next step:
- Quoting someone in an article or report
- Timestamped transcript: find the line, cite the time, jump back to check the wording
- Turning a talk into a blog post or handout
- Plain transcript: fewer timestamps to strip out
- Captions for your player, course platform or YouTube
- SRT or VTT
- Reading a video in a language you don't speak
- Pick a target language; the transcript download and subtitles follow it, and the result screen shows the original alongside
Pick a target language and one run gives both texts: Download transcript is the translation and Download original transcript is the wording as spoken. Subtitles come in one language per run, so captions in the spoken language plus a translated set means a second run; on a half-hour video that costs another 30 credits (3¢). Transcription or translation? explains how the two steps fit together.
Three things to check before uploading
- More than one audio track: camera files and some screen recordings carry several, and the one with speech isn't always the first. Videos with more than one audio track shows how to check and pick.
- Long music-only stretches: speech models can write words that nobody said during intros, outros and music beds. Trimming a two-minute title sequence is quicker than deleting phantom lines afterwards.
- Recordings longer than two hours: split at a natural pause, such as a break in a livestream or between conference sessions, rather than at an arbitrary time.
- Speakers far from the microphone: in a room recording, people far from the camera come out with more errors. If a lapel or desk microphone also recorded the session, transcribe that file instead of the camera's audio.
Cleaning up a machine transcript
- Search for names, product terms and numbers first; these are where recognition errors cluster.
- Use the timestamps to jump to doubtful passages instead of re-listening to the whole video.
- Add speaker names by hand. Transcripts don't label speakers, so a Q&A or interview needs names inserted where the voice changes.
- Decide how verbatim you need it. Speech models of this kind tend to smooth over ums and false starts; if your work needs true verbatim, add them by ear.
- Fix sentence breaks last, once the words are right.
The proofreading method for AI transcripts covers how to do this without rereading every line, and from transcript to article picks up where the clean transcript ends.
Worked examples
50 credits (5¢). The timestamped transcript becomes study notes where every heading carries a time to jump back to.
Transcript mode with English as the target: 12 credits (1.2¢). The English SRT doubles as captions if the demo is later shared internally.
Billed at the 5-credit minimum: 5 credits (0.5¢). If you have several short recordings, joining them into one file before upload avoids paying the minimum on each.
Where the text goes next
A transcript is rarely the final product. The common next steps each favour a different download:
- Searching a back catalogue: keep the plain transcripts in one folder next to the videos, named the same way, and ordinary desktop search finds the video where something was said.
- Cutting clips: scan the timestamped transcript for strong lines, note the times, and cut in your editor. Finding clips with a transcript shows the method.
- Publishing alongside the video: a cleaned transcript under an embedded video helps readers who prefer text and anyone who can't play sound. Formatting a raw transcript covers paragraphs, headings and speaker names.
- Sharing with colleagues who read another language: send the translated transcript and keep the original for checking quotes.
What a video transcript won't contain
- Who is speaking: there are no speaker labels.
- Words that appear only in the picture, like slide text or a caption already burned into the footage. Only speech is transcribed.
- Sound descriptions such as laughter or music cues.
- A summary: mydubly returns the full text and doesn't summarise it.
For the spoken language, see the pages for Hindi, where Hindi and English often mix in one sentence, Japanese, which is written without spaces between words, and Arabic. MP4 files covers the most common container, and the video transcription guide walks through a full run.
Transcribe speech in…
By file format
Learn more
Articles that go deeper on the technology and workflows behind this tool.
- Automatic speech recognition, explained from Audrey to Whisper
- OpenAI Whisper: how the open speech recognition model works
- Whisper large-v3-turbo: a faster Whisper with a four-layer decoder
- Word error rate explained, with a worked example
- Why Speech Recognition Struggles With Some Accents, and What Helps
- How One Speech Model Transcribes Many Languages, and Why Quality Varies
- How Speech Models Work Out Which Language Is Being Spoken
- A Faster Way to Proofread AI Transcripts Without Rereading Every Line
- When Whisper Writes Words Nobody Said: Hallucinations Explained
- From Transcript to Article: Turning a Video into a Blog Post
- Make a Lecture Archive Searchable with Transcripts
- From Recorded Lecture to Study Notes You Will Actually Use
- Quoting and Citing Video and Audio Sources with Timestamps
- Transcription or Translation? How They Differ and Fit Together
- How to Split Long Audio for Speech Recognition Without Cutting Words in Half
- Wind Noise in Outdoor Video: Why It Happens and What You Can Still Save
- Transcribing Recordings of Conferences, Ceremonies and Public Meetings
- Screen Recording with Audio: Microphone, System Sound and Clean Narration
- Formatting a Raw Transcript So People Can Actually Read It
- How to Turn a Video into an Email Newsletter
- Turning Product Walkthrough Videos into Help Center Articles
- Turning a Video Transcript into Text Posts for Social Media
- How to Publish Transcripts Alongside Your Videos and Podcasts
- Finding Short Clips in a Long Video Using Its Transcript
- Turning a Meeting Transcript into Useful Notes
- Turning Course Videos into Written Lessons and Handouts
- Turning Video Transcripts into Vocabulary Lists and Audio Flashcards
- Following University Lectures in a Language That Isn't Your First
- Using Authentic Video in ESL and EFL Lessons, from Clip Selection to Copyright
- Running an Accessible Webinar: Before, During and After
- Making Workplace Training Videos Work for Every Employee
- Descriptive Transcripts: Putting the Whole Video into Text
- Writing an Audio Description Script, Step by Step
- Clothing Rustle and Cable Thumps: Getting Clean Sound From a Lavalier
- How to Make Months of Meeting Recordings Searchable
- Video Metadata and Tagging for a Library You Can Search
- Using Transcripts in a Personal Knowledge Base Without Drowning in Them
- Knowledge Capture Interviews: Recording What Experts Know Before They Leave
- How to Fact-Check Claims Made in Videos and Podcasts
- How to Organize Family Videos So You Can Find Every Story
- Text-Based Video Editing: Cutting a Video by Editing Its Transcript
Frequently asked questions
Does it work for any video?
Any video your browser can read in MP4, MOV, WebM, MKV or M4V, up to 2 hours. DRM-protected purchases can't be read.
Can I convert a Zoom or Teams recording to text?
Yes, once it's a file on your device. Meeting recordings are usually saved or downloadable as MP4; download it first, then upload it here. A link to the recording won't work.
Which download should I open in Word or Google Docs?
The plain transcript or the timestamped .txt. SRT and VTT files open as text too, but they include cue numbers and timing lines you'd have to delete.
Why is there text in my transcript where only music plays?
Speech models sometimes produce words during music or silence. Delete those lines, and for future videos trim long music-only sections before uploading.