Research & academic workflows

How Long It Really Takes to Transcribe an Interview

Transcribing an interview by hand takes several times longer than the recording itself; a figure commonly quoted for an experienced typist with clear two-person audio is around four hours or more per hour of audio, and beginners or difficult recordings can take far longer. Checking an AI draft is usually much quicker, but it still takes at least the length of the recording, because you need to listen to all of it. The only reliable estimate for your project is one you measure yourself on a pilot interview.

7 min read · Updated

The short answer, with caveats

There is no single correct figure, because transcription time depends on the recording, the speakers, the level of detail you need and the person doing it. What can be said with confidence is the shape of the problem:

  • Manual transcription takes several times the recording length. Speech is much faster than typing, so you constantly pause, rewind and replay.
  • Correcting a machine draft is faster than typing from scratch for most clear recordings, but it is not instant. You still listen to everything, stop to fix errors, and add whatever the software left out.
  • Difficult audio, many speakers, strong accents, technical vocabulary and detailed notation all push times up, in both approaches.

Rules of thumb are useful for a first plan and dangerous as a final one. Treat any number in this article as a starting assumption to replace with your own measurement.

Manual transcription: what drives the time

A commonly quoted figure is that an experienced transcriber needs roughly four hours or more for each hour of clear, two-person interview audio. Newcomers often report that their first transcripts take considerably longer than that. These are rough, frequently repeated estimates rather than measured standards, and your own speed may differ widely.

The time goes on a cycle of listening to a few seconds, typing, rewinding to check, and continuing. Several factors change how fast that cycle runs:

  • Typing speed and accuracy. Touch typists have a large advantage.
  • Audio quality. Background noise, distance from the microphone, echo and low volume all mean more replays.
  • Number of speakers and overlap. Two people taking turns is far easier than a group talking over each other.
  • Accents, dialect and speaking pace relative to the transcriber's familiarity.
  • Vocabulary. Technical terms, names and acronyms need checking.
  • Level of detail. Clean verbatim is faster than true verbatim, and detailed notation for pauses and overlaps is much slower again.
  • Anonymization during transcription, if you replace names as you go.
  • Tools. Playback software with keyboard shortcuts, slowed playback, automatic rewind on pause, or a foot pedal saves a great deal of time.

Checking an AI draft: what drives the time

With an automatic draft, the work changes from typing to listening and correcting. At a minimum, a full correction pass takes as long as the recording, because you listen to all of it. In practice you also stop to fix words, add speaker labels and handle passages the draft got wrong.

How much longer than real time depends mainly on draft quality. A clear recording with standard vocabulary might need only occasional stops. A noisy recording with crosstalk, accented speech or specialist terms can need so many fixes that the draft saves little. The article comparing AI and human transcription explains why accuracy varies so much between recordings.

Typical extra work when correcting a draft:

  • Adding speaker labels, since many automatic tools do not provide them.
  • Correcting names, places, numbers and technical terms.
  • Restoring fillers and false starts if your style is verbatim; recognition models tend to drop them.
  • Checking passages where the draft has a gap, or a suspicious phrase in what should be silence.
  • Adding bracketed notes for laughter, long pauses or inaudible stretches.
  • Formatting to your project's conventions.

The guide to proofreading an AI transcript sets out an efficient order for those tasks. As an illustrative planning assumption only, not a measured figure, some researchers start by budgeting between the recording length and two to three times it for correction of clear interviews, then adjust once they have timed a pilot.

Factors that change the estimate either way

Clear two-person audio, close microphones
Fastest in both approaches
Group discussion with overlapping talk
Much slower; speaker attribution dominates the time
Noisy or outdoor recording
More replays manually; more errors to fix in a draft
Strong accents or dialect unfamiliar to the transcriber
Slower manually; recognition accuracy may also drop
Technical or local vocabulary
Time spent checking terms either way
True verbatim or detailed notation
Considerably slower, and drafts help less
Interview in a language the transcriber knows less well
Much slower; consider a fluent transcriber
Good playback tools and shortcuts
Noticeably faster manual work

Group recordings deserve special mention. Distinguishing voices and handling crosstalk can make a focus group take far longer than an interview of the same length; the article on focus group transcription covers why.

Measure your own speed with a pilot

The most useful thing you can do is time yourself:

  1. Pick a representative interview, not your best-quality recording.
  2. Transcribe or correct a ten-minute stretch in exactly the style you will use for the project, including speaker labels and any notation.
  3. Time it honestly, including breaks to check terms.
  4. Divide by ten to get minutes of work per minute of audio, and multiply by your total audio.
  5. Add a margin for harder recordings, fatigue and final checking.

If you plan to try both approaches, do the same ten minutes manually and from a draft. The comparison tells you whether automatic drafts are worth it for your recordings, and it gives your supervisor a defensible number.

Planning transcription into a thesis or project timeline

Transcription is often underestimated in research plans, partly because it feels like a mechanical step rather than research. It is better to plan it explicitly:

  • Total your expected audio: number of interviews times expected length.
  • Apply your measured rate, or a cautious rule of thumb until you have one.
  • Remember that transcription time is concentrated work; few people can transcribe accurately for a whole day. Plan in sessions of a few hours.
  • Transcribe as you go rather than after all interviews. Early transcripts improve later interviews and spread the load.
  • Include time for anonymization, member checking if you use it, and formatting for your analysis software.
  • Add time for checking quotes against the audio before submission.
Hypothetical: a master's dissertation with fifteen interviews

A student plans fifteen interviews of about fifty minutes each, roughly twelve and a half hours of audio. Using the commonly quoted figure of around four hours per audio hour as a rough guide, manual transcription would be in the region of fifty hours of focused work, and likely more as a beginner. She pilots one interview from an automatic draft and finds that correcting it, adding speaker labels and bracketed notes takes about twice the recording length. On that measured rate, the corrections come to around twenty-five hours. She schedules two correction sessions per week alongside her remaining interviews and keeps a buffer for two recordings made in a noisy café.

Common mistakes in estimating

  • Using someone else's figure without checking it against your own recordings.
  • Assuming an AI draft means transcription takes no time at all.
  • Piloting on your clearest recording, then hitting slower ones later.
  • Forgetting speaker labels, anonymization and formatting.
  • Leaving all transcription until data collection ends.
  • Underestimating fatigue; accuracy falls in long sessions.

How mydubly affects the timeline

mydubly can produce the draft you correct. You choose an interview recording from your device, up to two hours per file, and receive a plain transcript, a timestamped transcript with [m:ss] labels for jumping to passages, and SRT and VTT files. The spoken language is detected automatically, in any of 21 supported languages. Processing runs while the browser tab stays open; leaving or closing the page cancels the job. The audio to text page explains the inputs and outputs, and the interview transcription use case walks through a typical workflow.

It removes typing, not checking. You will still listen to every recording, add speaker labels, correct names and terms and add any notation. If your interviews will also be translated, translation turnaround is a separate question covered in the article on how long video translation takes. Check with your ethics board before uploading participant audio. On cost, fifteen fifty-minute interviews come to 750 credits (75¢).

Next step

Before you schedule anything else, run the ten-minute pilot described above on one real interview, in the style you will use for the project. Multiply out, add a margin, and put transcription sessions in your calendar alongside your interview dates rather than after them.

Frequently asked questions

How long does it take to transcribe a one-hour interview by hand?

A commonly quoted rough figure for an experienced transcriber with clear two-person audio is around four hours or more. Beginners, poor audio, many speakers, accents and verbatim detail can make it take much longer. Time yourself on a ten-minute sample to get a figure for your own recordings.

Is correcting an AI transcript faster than typing it myself?

For clear recordings it usually is, because you fix errors instead of typing every word. It still takes at least the length of the recording, since you need to listen to all of it. For very noisy or crosstalk-heavy audio, the saving can be small, so pilot both approaches if you are unsure.

Why does transcription take so much longer than the recording?

People speak far faster than most can type, so manual transcription means constant pausing, rewinding and replaying. Unclear words, overlaps, names and formatting add more stops. Even with a draft, each correction interrupts playback.

How should I plan transcription for a thesis?

Total your expected audio, multiply by a rate you have measured on a pilot, add a margin, and schedule transcription in regular sessions as interviews happen rather than all at the end. Include time for anonymization, formatting and checking quotes before submission.

Does slowing down playback help?

For manual transcription, slowed playback with keyboard shortcuts or a foot pedal often reduces rewinding and can save time overall. For correcting a good draft, normal or slightly faster speed is often enough, slowing down only for difficult passages.

How long does mydubly take to produce a transcript?

Processing time varies with the length of the file, so run a pilot recording to see what to expect. The browser tab must stay open while it runs, since leaving the page cancels the job. The bigger time cost in any case is your own correction pass, which you should plan for separately.