Education & research

Transcribing Oral History Interviews for the Archive

Oral history transcription turns recorded life-story interviews into text that researchers, families and future archivists can search, read and cite. Unlike a business transcript, it has to respect the narrator's consent and voice, document how decisions about verbatim and dialect were made, and survive in formats that will still open decades from now. AI transcription can produce the first draft quickly; the archival judgment around it remains human work.

7 min read · Updated

How oral history differs from other interview transcription

A journalist transcribes to find quotes, and the general interview transcription page covers that. A qualitative researcher transcribes to code themes, which the research interview transcription page covers. An oral historian transcribes to create a lasting document of one person's account, often read by people who will never hear the recording.

That changes the priorities. Faithfulness to the narrator's way of speaking matters more than tidiness. The narrator may review the transcript. And the file becomes an archival object in its own right, with metadata, rights and a preservation plan.

Verbatim decisions to make and write down

There is no single correct transcription style for oral history, but every project needs a documented one, decided before the first transcript.

Fillers and false starts
Keep all of them, keep those that show hesitation or emotion, or remove them. Many oral history programs lightly edit; linguistic projects keep everything.
Dialect and nonstandard grammar
Preserve the narrator's grammar and word choice. Avoid correcting dialect into the standard language, and avoid eye-dialect spellings that caricature speech.
Nonverbal events
Mark them consistently, for example [laughs], [long pause] or [crying].
Uncertain words
Use a convention such as [unclear: Brennan?] with a timestamp, so later readers can check the audio.
Interviewer speech
Usually transcribed in full and labeled with initials.

Whisper-style speech recognition leans toward clean output: it often drops fillers and repetitions and can regularize grammar slightly. That is a convenient start for a lightly edited style and a liability for strict verbatim. Either way, reviewers must listen against the recording rather than just read. The trade-offs are covered in verbatim vs clean verbatim transcription.

Older recordings and dialects

  • Digitize tapes properly before anything else. Archival guidance such as IASA TC-04 recommends uncompressed preservation masters at high resolution; keep the master untouched and make a separate access copy for transcription.
  • Hiss, hum, wow and flutter degrade recognition. Light cleanup on the access copy can help, but heavy noise reduction can add artifacts that make things worse, so test on a short section; see speech recognition and background noise.
  • Recognition models are trained mostly on contemporary speech, so older pronunciations, strong regional dialects and switching between languages can raise error rates, as explained in speech recognition and accents.
  • Place names, family names and local terms will be the most common errors. Build a project word list from research notes and give it to every reviewer.
  • Narrators often speak more quietly than interviewers, who tend to sit closer to the microphone, so listen for passages where the narrator is barely audible.

A technical point saves storage and confusion: speech recognition typically runs on 16 kHz mono audio, and mydubly decodes every file to 16 kHz mono before recognition. A 96 kHz master therefore gives no recognition advantage over a good access copy. Keep high resolution for preservation, not for transcription.

A workflow for an oral history project

  1. Digitize each recording, creating a preservation master and an access copy, and record checksums so you can verify the files later.
  2. Name files with the collection identifier, such as OH-2026-014_session1.wav, and log the interview's metadata at the same time.
  3. Generate a draft timestamped transcript from the access copy.
  4. Review against the audio at full speed, correcting words, adding speaker initials and marking uncertain passages with timestamps.
  5. Apply the project style guide for fillers, dialect and nonverbal marks.
  6. Send the transcript for narrator review where the project offers it, and log the changes made.
  7. Proofread once more, add a metadata header and export to archival formats.
  8. Build a timestamped index of topics so users can jump to the parts of the audio they need. Dedicated tools such as OHMS, the Oral History Metadata Synchronizer, exist for this.
A two-session interview, costed

Suppose a community history project records one narrator over two sessions of 95 and 140 minutes. The second session exceeds the 2-hour file limit, so it is split at a break into two 70-minute parts. Drafts for all 235 minutes cost 235 credits, or 23.5¢. The review, which means listening, fixing family names and dialect terms and adding speaker initials, is where the project's hours actually go.

Metadata that should travel with every transcript

Collection and item identifier
Links the transcript to its recording and release form.
Narrator and interviewer
Names, or pseudonyms where access is restricted, and their roles.
Date, place and duration
For each session.
Languages
Including dialect notes and any switching between languages.
Rights and access
Copyright holder, release terms, restrictions and when they expire.
Transcription method
Draft tool, style guide version, reviewers and review dates.
Recording source
The original carrier and how it was digitized.

Recording the transcription method matters more than it seems. A researcher thirty years from now should be able to tell that a draft was machine-generated and then human-reviewed, and which style rules applied. Dublin Core-based schemas used by many archives have fields for most of this.

Formats that will still open in decades

  • Master transcript as plain UTF-8 text, with a PDF/A copy as the formatted reading version. Avoid making a proprietary word processor file the only copy.
  • Timed text kept alongside: the timestamped transcript and SRT or VTT files are plain text that future players can sync with the audio.
  • Audio preservation masters as WAV or FLAC, with MP3 or M4A access copies for listening.
  • Several copies in different places, checked against their checksums on a schedule.

What AI drafts can and can't do for oral history

AI transcription is good at the part of the job that used to block small projects: getting from hours of audio to editable text.

  • It turns typing into review, which makes volunteer-run projects with tape backlogs feasible.
  • Timestamps arrive with the draft, which makes indexing and later fact-checking easier.
  • Translated transcripts can open interviews in other languages to a wider audience, provided someone fluent checks them.

It cannot replace the judgment that makes a transcript archival.

  • There are no speaker labels, so interviewer and narrator run together until a reviewer adds initials.
  • Output leans clean, which conflicts with strict verbatim styles.
  • Dialect, worn recordings and proper nouns remain error-prone.
  • Sending audio to any external service may conflict with a narrator's consent or a community's protocols, so check before uploading.

Using mydubly for oral history drafts

mydubly accepts audio files (WAV, FLAC, MP3, M4A, AAC, OGG) and video files from your device, up to 2 hours each, and returns a timestamped transcript plus SRT and VTT files. The spoken language is detected automatically across 21 languages, and a translation can be requested. Transcripts cost 1 credit per minute with a 5-credit minimum per file, and credits do not expire, which suits projects that work through tapes over many months. FLAC access copies work directly; see FLAC to text.

For consent forms, describe the processing accurately: the browser extracts the audio track and uploads it in chunks, and uploaded audio and results are deleted within 30 minutes of a job finishing, while unfinished jobs expire after 24 hours. For video interviews the picture never leaves the device. Transcripts do not label speakers. The audio to text page has the details.

Begin with one interview

Write a one-page style guide, choose an interview with good audio, generate a draft with audio to text and time the review. That figure tells you what the whole collection will take and whether you need more reviewers. When researchers later need to cite passages from your collection, point them to how to cite a video with timestamps.

Frequently asked questions

Should an oral history transcript correct the narrator's grammar?

Generally not. The transcript documents how the narrator spoke, so grammar and word choice are preserved, while spelling follows standard conventions. Whatever you decide, write it in the style guide so every transcript in the collection is consistent.

Do we need the narrator's permission to use automatic transcription?

Your consent and release forms should describe how the recording will be processed, including any external transcription service. Telling narrators up front avoids surprises, and some community protocols require it. Check your archive's or institution's rules.

What sample rate should we digitize tapes at?

Follow preservation guidance such as IASA TC-04 or your archive's own standard for the master. For transcription the resolution matters much less, because recognition typically runs on 16 kHz mono audio, so a good access copy is enough.

How do we transcribe an interview that switches between two languages?

The spoken language is detected automatically from the audio, so passages in different languages can be handled inconsistently. Review those stretches closely, and consider the workarounds in translating mixed-language videos.

Is a VTT or SRT file useful for an audio-only oral history?

Yes. Some archive platforms use timed text to show the transcript scrolling in sync with the audio, and the files are plain text that preserve well. Keep them alongside the master transcript even if you do not use them yet.