About MP4 files
MP4 is the most common video container: phones, screen recorders, video editors and meeting apps all export it, usually with H.264 video and AAC audio.
Common sources: phone cameras, Zoom, Teams and Meet recordings, Premiere, DaVinci Resolve and CapCut exports, screen recorders.
Every modern browser can read MP4, so it is the most reliable format to start with.
MP4 to text
If your MP4 has several audio tracks (for example a separate commentary track), only one is used — export the file with your main speech as the only or main audio track.
A team lead downloads a 48-minute Teams recording as MP4 and turns it into a transcript to write up decisions.
Cost: 48 credits (4.8¢).
Translating a MP4
MP4 in, MP4 out: the translated video keeps your original picture and swaps in the translated voice track.
A small business translates a 6-minute MP4 product video into French and uploads the translated MP4 to its Canadian site.
Cost: 300 credits (30¢) with a translated voice, or 6 credits (0.6¢) for translated text only. Switch Full translation output on in the tool above for a voice track.
What an MP4 actually holds
An MP4 is built from boxes. An index box (moov) tells a player where every video frame and audio sample sits, and a data box (mdat) holds the compressed streams themselves. For a transcript only one stream matters: the audio, which is usually AAC and occasionally Opus or Dolby AC-3 in files from TV tuners and some cameras. The video stream, whether H.264, HEVC or AV1, is never decoded for text, which is why a 4K recording costs the same per minute as a 480p one.
The index box is also the weak point. Recorders write it last, so a recording cut off by a crash, a full card or a flat battery often has no index and won't open anywhere. Repair tools can sometimes rebuild it using a healthy file from the same device as a reference. How MP4 files are structured explains the boxes in more detail.
Where MP4s come from, and the quirk each source brings
- Zoom local recording
- One mixed audio track. Zoom also saves an audio-only M4A beside the video that carries the same sound and is much smaller to handle.
- Teams or Meet recording
- Teams stores recordings in OneDrive or SharePoint and Meet in Google Drive. Download the MP4 to your device first, because a share link can't be pasted in.
- Phone camera
- Often variable frame rate and HEVC video, with stereo sound from two microphones a few centimetres apart.
- Premiere, Resolve or CapCut export
- Exactly what you mixed. A loud music bed in the export competes with the speech in the transcript too.
- Screen recorder
- Microphone and system sound may land on separate tracks, or system sound may be missing entirely.
Three things to check before uploading
Multiple audio tracks. OBS, some screen recorders and multicam workflows write two or more audio tracks into one MP4, for example the microphone on one and game or desktop sound on another. Only one track is used, so open the file in VLC (Audio → Audio Track) or a media inspector such as MediaInfo and make sure the speech is the only track, or at least the main one. If it isn't, re-export with a single mixed track. Multiple audio tracks in video files walks through the options.
Variable frame rate. Phones and screen recorders vary the frame rate to save space. A transcript is unaffected, because its timing comes from the audio. It matters if you plan to dub: play the original first, and if lips already drift away from the voice towards the end, convert to constant frame rate in your editor before translating. Variable frame rate explained shows how to spot it.
Audio the browser can't decode. AAC and Opus play everywhere, but Dolby AC-3 and E-AC-3 aren't supported by every browser. If a file loads silently or not at all, re-export the sound as AAC; the picture can be copied across untouched.
What an MP4 job hands back
Transcript mode gives you a timestamped .txt, SRT and VTT subtitles and a plain transcript; pick a target language and the subtitles and plain text arrive translated. Dub mode gives you the translated video: an MP4 with your original picture and the new voice, or an MKV when the original video codec can't be placed in an MP4, since the merge happens in your browser. Alongside it come the translated audio as M4A, the original-language transcript, and SRT and VTT that follow what the voice actually says, including lines shortened to fit the timing.
Pricing a recurring meeting
A 75-minute Teams recording costs 75 credits (7.5¢) for a transcript, timestamps and subtitles. A second run with Polish picked, for the Warsaw office, is another 75 credits (7.5¢). Dubbing only the 20-minute CEO segment, cut out in an editor first, costs 1,000 credits ($1.00).
For meeting workflows, meeting transcription and turning a recording into notes pick up where the transcript stops. If your video came off an iPhone as a MOV, the MOV page covers the Apple-specific quirks.
Using the tool on this page
- The spoken language is detected automatically; you select the output language — the spoken language itself for an untranslated transcript, or another language for a translation. Files can be up to 2 hours long.
- Only the audio is sent for processing — the video stays on your device — and audio and text are deleted within 30 minutes of delivery (how files are protected).
- You can download a timestamped transcript, subtitles (SRT and VTT) and plain text. Every output, step and limit is explained on the video to text page.
Frequently asked questions
Do I need to convert my MP4 to MP3 before transcribing it?
No. Drop the MP4 in directly — your browser pulls out the audio itself, and the video never leaves your device.
Will the translated file also be an MP4?
Yes. Your browser combines the original picture with the translated audio and saves an MP4. On devices that can't handle a very large file, you get the translated audio track to download instead.
My MP4 is several gigabytes. Is that a problem?
Size mostly affects how long your browser takes to open the file. The sound is pulled out on your device and only that is sent, so a 6 GB file doesn't mean a 6 GB upload. What's limited (2 hours) and billed is length. On a phone or an older laptop a very large file can take a while to read, so use a computer for multi-gigabyte recordings.
Why does my transcript start with a line nobody said?
Speech recognition models can invent a short phrase over silence or music, for example in the minute before a meeting starts. Trim dead air from the start and end of the MP4, or just delete those lines. Whisper hallucinations explains why it happens.
Can I transcribe just one section of a long MP4?
Trim it first. Billing follows the length of the file you upload, so cutting a 90-minute recording down to the 15 minutes you need saves credits. A cut using stream copy in an editor or ffmpeg keeps the original quality.