Transcription · English

English Transcription: Video & Audio to Text

Get a English transcript of any video or recording spoken in English — plain text, timestamped, or as subtitles. Select English as the language to keep it untranslated, or another language to have it translated.

01 — Upload

Your video

Full translation outputTimestamped transcript only — no video or audio in results.

The video file stays in this browser. Language is required: pick the spoken language for a plain transcript, or another language to translate it.

02 — Result

Your transcript

Your timestamped transcript will appear here.

Waiting for upload

25 languages available

How English speech transcribes

English is the language speech-recognition models have seen most, so clear recordings come out with few errors. Background noise and several people talking at once cause more mistakes than accents do.

What a English transcript looks like

English transcripts come back punctuated and in sentence case, with numbers usually written as digits. Filler words like 'um' are mostly dropped, which makes the text easier to read but means it isn't a strict verbatim record — add them back if you need legal-style verbatim.

Examples

60-minute video

A product manager transcribes a 60-minute customer call exported from the meeting app as MP4 to pull exact quotes for a roadmap review.

25-minute recording

A journalist turns a 25-minute phone interview saved as M4A into searchable text before writing the story.

Costs: 60 credits (6¢) and 25 credits (2.5¢), including a translated transcript if you select another language. Audio files go in the audio to text workspace.

When English is everyone's second language

A large share of English-language video is recorded by people for whom English is a second or third language: international team calls, supplier demos, conference talks, university seminars. A single clear accent is rarely the problem. What trips recognition up is a word pronounced in an unexpected way that also happens to be a real English word: "ship" written as "sheep", "thirty" as "thirteen", a product name turned into a common noun. The sentence still reads as grammatical English, so a quick skim won't catch it.

For these recordings, proofread by meaning rather than by spelling. Read each sentence asking whether this speaker would plausibly have said it, and play the timestamp for anything that sounds off. Numbers deserve a second listen in every recording with non-native speakers, because the teen/ty pairs are the classic confusion. The article on speech recognition and accents explains why some accents produce more substitutions than others.

Acronyms, product names and US or UK spelling

English business and technical video is dense with terms a general model has never seen: internal project codes, drug names, ticker symbols, the name of your startup. Expect some of these to come back as the nearest ordinary words or as a plausible but wrong acronym. Before editing line by line, list the ten or so terms the speaker repeats most, then fix each one with find-and-replace across the whole file.

Medical webinar terms list

Write down drug names and lab terms such as metformin, HbA1c and GLP-1 before reading. One find-and-replace per term catches every variant at once, and the line-by-line read can then focus on meaning instead of spelling.

Spelling convention is a separate decision. Check which variety the transcript uses for pairs like colour/color, organise/organize and programme/program, and whether it matches your publication. If it doesn't, normalise it in one pass at the end; a quote that mixes conventions looks careless in print.

How far from verbatim the text is

The output reads as clean English, which suits captions, blog posts and notes. Researchers, court reporters and linguists sometimes need every false start, repetition and hesitation instead. Beyond the dropped fillers mentioned above, look for self-corrections ("we shipped, sorry, we'll ship in May") that may be smoothed into one version, and for words spoken over each other in a discussion, which may be missing. If your field requires a strict record, treat the transcript as a first draft and listen through the whole file; verbatim versus clean verbatim sets out which style each kind of project needs.

Transcripts don't label speakers, so in an interview or panel you add names yourself. Speaker changes often fall at pauses, so scanning the timestamped transcript for gaps between lines is a quick way to place them.

Clips, songs and music beds

English-language video often borrows audio: a news clip inside a commentary, a song under a montage, a movie quote in a lecture. Lyrics and clip dialogue may be transcribed as if the presenter said them, and long stretches of music with no speech can occasionally produce a line of text nobody spoke. Before publishing captions, play the parts of the timeline where music is loud and delete anything that doesn't belong. Whisper hallucinations explains why silence and music cause invented text.

Costed: a 45-minute all-hands with speakers from five countries

Recording
45-minute MP4 exported from the meeting app
Mode
Transcript mode, English selected (no translation)
Cost
45 credits (4.5¢)
Downloads used
Timestamped transcript for the minutes, SRT for the internal video portal
Review focus
Figures, project code names, any sentence that reads oddly

If colleagues in Madrid and Seoul want the text in their own language, run the file again with Spanish or Korean as the target. Each extra language is a separate run of 45 credits, and that run's subtitles and plain transcript come back translated, with the original shown alongside on the result screen.

Captions from an English transcript

For same-language captions, select English as the language, so the text comes back untranslated, and download the SRT or VTT. English cues are easy to adjust in a text editor; the most common fixes are joining a cue that splits a name across two lines and trimming cues where a fast talker produces more text than a viewer can read in time. The SRT generator page shows the file structure, and meeting transcription walks through the workflow for recorded calls.

Using the tool on this page

  • The spoken language is detected automatically; you select the output language — the spoken language itself for an untranslated transcript, or another language for a translation. Files can be up to 2 hours long.
  • Only the audio is sent for processing — the video stays on your device — and audio and text are deleted within 30 minutes of delivery (how files are protected).
  • You can download a timestamped transcript, subtitles (SRT and VTT) and plain text. Every output, step and limit is explained on the video to text page.

Other languages

Frequently asked questions

Does it handle different English accents?

Yes — American, British, Indian, Australian and other accents generally transcribe well. Heavy background noise or overlapping speakers cause more errors than accents do.

Does the transcript use American or British spelling?

Check a few common words such as color/colour in your own output rather than assuming the speaker's accent decides it. Whichever convention appears, one find-and-replace pass per word pattern brings the transcript in line with your style guide.

Can I transcribe an English video and get the text in another language?

Yes. Pick one of the other 20 languages as the target. The result screen shows the English and the translation side by side, and the subtitles and plain transcript you download are in the target language. It is still 1 credit per minute.

Is an English transcript good enough to quote from?

For quotes, always play back the exact passage before publishing. Misheard words in English usually produce a different real word, so a quote can read naturally and still be wrong. The timestamped transcript takes you straight to the spot.