Comparisons

Transcription or Translation? How They Differ and Fit Together

Transcription turns spoken words into written text in the same language; translation renders that meaning in a different language. In transcription vs translation, the first is about accuracy to what was said, the second about accuracy to what was meant. For audio and video they usually run in sequence: the speech is transcribed first, and the transcript is what gets translated.

6 min read · Updated

The difference at a glance

The two words get used interchangeably in requests and quotes, which leads to surprises when the deliverable arrives. Here is how they compare for recorded speech.

Input and output
Transcription: speech in, text in the same language out. Translation: text (or speech) in one language in, text in another language out.
What "correct" means
Transcription: the words match what was said. Translation: the meaning, tone and intent match the source.
Typical files
Transcription: a timestamped transcript and same-language subtitle files. Translation: translated subtitles, a translated transcript, or a dubbed voice track.
Main failure
Transcription: misheard words, names and numbers. Translation: wrong meaning, wrong register, terms rendered inconsistently.
Who can check it
Transcription: anyone fluent in the spoken language. Translation: someone fluent in both languages.

What transcription produces

A transcript is a written record of speech. Machine transcripts from modern speech recognition models, such as OpenAI's Whisper, come back as segments of text with start and end timestamps, which is what makes them useful beyond reading: you can jump to a moment, quote it precisely or turn the segments into subtitle cues.

Transcription is also where decisions about style happen. Whether to keep "um" and false starts, how to write numbers and how to handle crosstalk are all transcription questions, and they affect everything built on top of the text.

What translation adds

Translation takes that text and re-expresses it in another language. Good translation is not word replacement. Word order changes, idioms need equivalents, formality has to suit the audience, and some languages need information the source never stated, such as a speaker's gender or whether "you" is singular or plural.

For video, translation can be delivered in different forms: subtitles the viewer reads, a transcript in the target language, or a voice track that speaks the translation. The form you need depends on how the audience will watch, which is covered in subtitles vs dubbing.

How the two chain together

In speech translation the transcript is the foundation, so its errors flow downstream. A translation engine cannot know that a word was misheard; it faithfully translates the wrong word.

Before and after: a misheard word, translated

Spoken: "Let's review the Q3 numbers before the board call." Transcribed with an error: "Let's review the cute three numbers before the board call." Translated into Spanish from that transcript: "Revisemos los tres números lindos antes de la llamada con la junta." The translation is fluent and grammatical, and completely wrong, because the error happened one step earlier.

This is why reviewing the source-language transcript matters even when the translation is the deliverable. Names, product terms, acronyms and figures are the usual weak points, and they are much easier to spot in the original language than after translation.

When a transcript alone is enough

Plenty of jobs never need translation. A transcript is the right deliverable when:

  • The audience speaks the same language and wants to search, skim or quote the recording, as with meeting transcription or interviews.
  • You need same-language captions for accessibility or for viewers watching with the sound off.
  • You are repurposing a talk into an article, notes or show notes.
  • A researcher needs to code and analyze what participants said in their own words.

When you need translation as well

Translation earns its place when the people who need the content do not speak its language: a course offered to a new market, a training video for staff abroad, a podcast reaching listeners in another country, or a family video for relatives. It is also useful internally, when a reviewer needs to understand a recording in a language they do not speak before deciding what to do with it.

Where each falls short

Neither process is a complete solution on its own, and machine versions of both have predictable gaps.

  • Transcription limits: recognition struggles with heavy noise, overlapping speakers, strong accents underrepresented in training data and unusual names. Machine transcripts often do not label speakers, and mydubly's do not.
  • Translation limits: humor, wordplay and culturally specific references often translate flatly. Context beyond a sentence or a short passage can be lost, so pronouns and formality may drift.
  • Combined limits: errors compound. A small transcription slip can become a confident mistranslation, and two languages are needed to catch it.
  • Neither touches the picture: text shown on screen is not part of the speech, so it is neither transcribed nor translated.

Transcription and translation in mydubly, with costs

mydubly offers both as separate modes, and the spoken language is detected automatically, so you only choose the output language: the spoken language for untranslated text, or another language for a translation.

Transcript mode
Timestamped transcript plus SRT and VTT subtitles in the spoken language, optionally translated, with no voice. 1 credit per minute, minimum 5 credits per file.
Full translation
Translated MP4 with a new AI voice track, the translated audio file, SRT and VTT subtitles, and transcripts in both languages. 50 credits per minute, minimum 2 minutes per file.
Credits
$1 buys 1,000 credits, the minimum top-up is $1, and credits do not expire.

For example, transcribing a 30-minute interview costs 30 credits, which is 3 cents. A 3-minute voice memo is billed at the 5-credit minimum, half a cent. A full translation of the same 30-minute interview, with a dubbed voice track, costs 1,500 credits, or $1.50. You can start with video to text for transcripts, or the video translator for the full output; both work in 21 languages.

A sensible order of work

If you are not sure which you need, work in this order:

  1. Run a short, representative clip in transcript mode and read the result against the audio. This shows how well the speech is recognized before you spend more.
  2. If recognition is weak, improve the source before going further; improving transcription accuracy covers the usual fixes.
  3. Decide who the audience is. Same language: stop at the transcript and subtitles. Other languages: continue.
  4. Run the translation, choosing subtitles or a dubbed voice according to how people will watch.
  5. Review the source transcript and the translated transcript side by side, starting with names, numbers and key terms.

Where to go next

If a transcript is all you need, upload a file on video to text and download the timestamped text and subtitle files. For more on judging the translated side, see how accurate AI video translation is.

Frequently asked questions

Is a translated transcript the same as a translation of the video?

It covers the same words, but as a document rather than something viewers watch. A translated video adds the translation as subtitles or a voice track timed to the picture, while a translated transcript is for reading, searching and review.

Can translation skip the transcription step?

Some research systems translate speech directly, but most practical pipelines, including mydubly's, transcribe first and translate the text. Keeping the transcript makes errors traceable, because you can see whether a mistake came from hearing or from translating.

Who should review a machine translation of a recording?

Ideally someone fluent in both languages, with the source transcript open next to the translation. A reviewer who only reads the target language can catch awkward phrasing but not meaning that drifted from the original.

Do I need to tell mydubly which language is spoken?

No. The spoken language is detected automatically from the audio. You only pick the output language: the spoken language itself for an untranslated transcript, or another language for a translation.

Why is a dubbed translation so much more expensive than a transcript?

Generating and timing a new voice for every line is far more work than producing text. In mydubly, transcripts cost 1 credit per minute and the full translated output with voice costs 50 credits per minute.