Video to text

Video to Text Converter

Get every word of a video as text: a clean transcript, a timestamped version and subtitle files — in the spoken language, or translated. Your video stays on your device.

01 — Upload

Your video

Full translation outputTimestamped transcript only — no video or audio in results.

The video file stays in this browser. Language is required: pick the spoken language for a plain transcript, or another language to translate it.

02 — Result

Your transcript

Your timestamped transcript will appear here.

Waiting for upload

25 languages available

What people turn videos into text for

Meetings
Minutes and action items from Zoom, Teams and Meet recordings — see meeting transcription.
Lectures
Searchable notes from recorded classes — see lecture transcription.
Content
Blog posts, show notes and quotes from videos you've published.
Accessibility
Text and captions for viewers who can't hear the audio.

Priced for long recordings

Transcription costs 1 credit per minute — 60 credits (6¢) for an hour, 120 credits (12¢) for two. Adding a translated transcript costs nothing extra.

How it works

  1. Drop in a video. Your browser reads it and pulls out only the audio — the video file never leaves your device.
  2. The audio is processed in 30-second chunks: speech is transcribed and the spoken language is detected automatically.
  3. Download the results. Audio and text on our servers are deleted after delivery, within 30 minutes at most.

What you get

Transcript
Plain text (.txt) of what was said
Timestamped transcript
Every line prefixed with its time, e.g. [12:04]
Subtitles
SRT and VTT, in the language you select (select the spoken language for untranslated subtitles)

Supported files

  • Video: MP4, WebM, MOV, MKV, M4V — whatever your browser can read.
  • Length: up to 2 hours per file.
  • Languages: speech in 21 languages is recognised automatically; translation and voices cover the same 21.

Privacy

Your video is never uploaded — only its audio is, over HTTPS in small chunks. Audio, transcripts and voice files are deleted after delivery, within 30 minutes at most, and are never used to train models. Your account keeps only the file name, length, languages and credits used. Details in the privacy policy.

Good to know

  • No speaker labels — transcripts don't say who is speaking.
  • Files only: no links from YouTube or social apps, and no live audio.
  • Subtitles come as SRT/VTT files rather than burned into the picture.

Picking the right download for the job

Transcript mode gives four files, and each suits a different next step:

Quoting someone in an article or report
Timestamped transcript: find the line, cite the time, jump back to check the wording
Turning a talk into a blog post or handout
Plain transcript: fewer timestamps to strip out
Captions for your player, course platform or YouTube
SRT or VTT
Reading a video in a language you don't speak
Pick a target language; the transcript download and subtitles follow it, and the result screen shows the original alongside

Pick a target language and one run gives both texts: Download transcript is the translation and Download original transcript is the wording as spoken. Subtitles come in one language per run, so captions in the spoken language plus a translated set means a second run; on a half-hour video that costs another 30 credits (3¢). Transcription or translation? explains how the two steps fit together.

Three things to check before uploading

  • More than one audio track: camera files and some screen recordings carry several, and the one with speech isn't always the first. Videos with more than one audio track shows how to check and pick.
  • Long music-only stretches: speech models can write words that nobody said during intros, outros and music beds. Trimming a two-minute title sequence is quicker than deleting phantom lines afterwards.
  • Recordings longer than two hours: split at a natural pause, such as a break in a livestream or between conference sessions, rather than at an arbitrary time.
  • Speakers far from the microphone: in a room recording, people far from the camera come out with more errors. If a lapel or desk microphone also recorded the session, transcribe that file instead of the camera's audio.

Cleaning up a machine transcript

  1. Search for names, product terms and numbers first; these are where recognition errors cluster.
  2. Use the timestamps to jump to doubtful passages instead of re-listening to the whole video.
  3. Add speaker names by hand. Transcripts don't label speakers, so a Q&A or interview needs names inserted where the voice changes.
  4. Decide how verbatim you need it. Speech models of this kind tend to smooth over ums and false starts; if your work needs true verbatim, add them by ear.
  5. Fix sentence breaks last, once the words are right.

The proofreading method for AI transcripts covers how to do this without rereading every line, and from transcript to article picks up where the clean transcript ends.

Worked examples

50-minute recorded lecture in English

50 credits (5¢). The timestamped transcript becomes study notes where every heading carries a time to jump back to.

12-minute product demo in Korean for a team that reads English

Transcript mode with English as the target: 12 credits (1.2¢). The English SRT doubles as captions if the demo is later shared internally.

4-minute screen recording

Billed at the 5-credit minimum: 5 credits (0.5¢). If you have several short recordings, joining them into one file before upload avoids paying the minimum on each.

Where the text goes next

A transcript is rarely the final product. The common next steps each favour a different download:

  • Searching a back catalogue: keep the plain transcripts in one folder next to the videos, named the same way, and ordinary desktop search finds the video where something was said.
  • Cutting clips: scan the timestamped transcript for strong lines, note the times, and cut in your editor. Finding clips with a transcript shows the method.
  • Publishing alongside the video: a cleaned transcript under an embedded video helps readers who prefer text and anyone who can't play sound. Formatting a raw transcript covers paragraphs, headings and speaker names.
  • Sharing with colleagues who read another language: send the translated transcript and keep the original for checking quotes.

What a video transcript won't contain

  • Who is speaking: there are no speaker labels.
  • Words that appear only in the picture, like slide text or a caption already burned into the footage. Only speech is transcribed.
  • Sound descriptions such as laughter or music cues.
  • A summary: mydubly returns the full text and doesn't summarise it.

For the spoken language, see the pages for Hindi, where Hindi and English often mix in one sentence, Japanese, which is written without spaces between words, and Arabic. MP4 files covers the most common container, and the video transcription guide walks through a full run.

Transcribe speech in…

By file format

Learn more

Articles that go deeper on the technology and workflows behind this tool.

Frequently asked questions

Does it work for any video?

Any video your browser can read in MP4, MOV, WebM, MKV or M4V, up to 2 hours. DRM-protected purchases can't be read.

Can I convert a Zoom or Teams recording to text?

Yes, once it's a file on your device. Meeting recordings are usually saved or downloadable as MP4; download it first, then upload it here. A link to the recording won't work.

Which download should I open in Word or Google Docs?

The plain transcript or the timestamped .txt. SRT and VTT files open as text too, but they include cue numbers and timing lines you'd have to delete.

Why is there text in my transcript where only music plays?

Speech models sometimes produce words during music or silence. Delete those lines, and for future videos trim long music-only sections before uploading.