What a raw machine transcript gives you
Automatic transcripts arrive in one of two shapes. A plain-text transcript is the spoken words as continuous text, punctuated but with no structure. A timestamped transcript puts each recognized speech segment on its own line with its start time in brackets, such as [4:10] for four minutes ten seconds, or [1:02:37] past the hour. Neither says who is speaking, neither has headings, and the segment breaks follow pauses in speech rather than meaning.
That raw output is excellent material and a poor document. A reader cannot tell the interviewer from the guest, cannot skim to the part about pricing, and gets a timestamp every few seconds whether they need one or not. Formatting fixes all three without changing a word of what was said, which is a separate decision covered in verbatim vs clean verbatim transcription.
Decide what the transcript is for
The purpose decides almost every formatting choice, so settle it before you touch the text.
- Published alongside a video or podcast
- Reader-friendly paragraphs, speaker names, sparse timestamps or none, headings for long episodes.
- Research interview or focus group
- Speaker codes instead of names, timestamps at every turn, line or paragraph numbers for coding, notation for pauses and unclear words.
- Journalism and quoting
- Speaker names, frequent timestamps so every quote can be checked against the recording, no tidying of the words being quoted.
- Internal meeting record
- Names, headings per agenda item, timestamps per section; often condensed into notes rather than kept whole.
- Legal or formal proceedings
- Formats are set by the court, agency or reporting standard involved. Follow their template; a general guide like this one does not replace it.
Paragraphs and speaker turns
Start a new paragraph every time the speaker changes. Within one long answer, start a new paragraph when the topic shifts, or roughly every four to six sentences, so the page never becomes a wall of text. Ignore the machine's segment boundaries: a segment can end mid-sentence because the speaker paused, and two short segments often belong in one sentence.
Spoken language runs on. Where a speaker strings three thoughts together with "and so", you may split them into sentences in a clean transcript, as long as the meaning and the words stay the same. In a verbatim transcript, leave the run-on as spoken.
Adding speaker names when the transcript has none
Machine transcripts from most tools, mydubly included, do not label speakers; the reasons are explained in what speaker diarization is. You add names during review, listening at each change of voice. Pick one convention and apply it throughout:
- Full name at the first turn, then a short form: MARIA LOPEZ the first time, LOPEZ or ML afterward.
- Role labels where names should not appear: INTERVIEWER and PARTICIPANT, or Q and A for a straightforward interview.
- Codes for anonymized research: P1, P2, P3 for participants and M for the moderator, with the key stored separately from the transcript.
- Put the label at the start of the paragraph, followed by a colon, in capitals or bold, so it can be scanned and searched.
Searching for a question mark is a quick way to find likely speaker changes in an interview, because a question is usually followed by a different voice.
Timestamps: how many and where
A timestamp on every machine segment is too many for reading and too few for nothing; choose an interval that suits the purpose.
- At each speaker turn: the standard for interviews, research and journalism.
- At each paragraph: good for long monologues such as lectures and talks.
- At fixed intervals, every 30 seconds or every minute: useful when the transcript is mainly a navigation aid.
- At section headings only: enough for published podcast transcripts where readers rarely need the exact second.
To thin out a machine transcript, keep the timestamp of the first segment in each paragraph and delete the rest. Write timestamps in one format throughout. Hours, minutes and seconds with leading zeros, such as [00:04:10], sorts and aligns well; [4:10] is friendlier for short recordings. If timestamps will link back to an online video, match the format the platform understands.
Marking what you could not hear
Every transcript has moments that cannot be transcribed with confidence. Mark them consistently, in square brackets so they never look like speech:
- [inaudible 00:12:31] where words cannot be made out, with the time so someone can check.
- [unclear: Halvorsen?] for a best guess you are not sure of.
- [crosstalk] where two people speak at once and neither can be separated.
- [laughs], [pause], [music] for non-speech events, if the purpose needs them. Published transcripts usually keep few of these; research transcripts may keep many.
- Em dashes for interruptions and cut-off words, and ellipses for trailing off, if your style uses them.
Machine transcription tends to resolve unclear audio into confident-looking words rather than mark it, so these notes are added by a person listening at the doubtful spots.
A header block and headings
Put the recording's details at the top, so the transcript makes sense on its own months later:
- Title, date of recording and length.
- Participants with roles, or the code key reference for anonymized transcripts.
- Source file name, so the transcript can always be matched to its recording.
- Transcription method and style, for example "machine transcript, reviewed and lightly cleaned by J. Park".
- Any notes on editing, such as removed personal details or skipped sections.
For anything over about 20 minutes, add headings at topic changes. They double as a table of contents and make it easy to turn the transcript into chapters or show notes later.
Raw timestamped lines: "[4:10] so the main thing we found was that" then "[4:12] people didn't read the instructions at all did you expect that" then "[4:16] honestly no we assumed they would at least skim them". Formatted: "LOPEZ [00:04:10]: So the main thing we found was that people didn't read the instructions at all." New paragraph: "INTERVIEWER [00:04:14]: Did you expect that?" New paragraph: "LOPEZ [00:04:16]: Honestly, no. We assumed they would at least skim them." The words are the same; the speaker change was hidden in the middle of a machine segment.
From raw output to finished document
- Download the timestamped transcript, and the plain text if you prefer to work from it.
- Fill in the header block.
- Listen through at normal or slightly faster speed, inserting speaker labels and paragraph breaks at each voice change.
- Thin the timestamps to your chosen interval.
- Correct names, numbers and misheard terms; a review method is in proofreading an AI transcript.
- Add bracketed notes where audio is unclear, and headings at topic changes.
- Apply your house style for numbers, spelling variants and filler words.
- Export in the format readers need, and keep an unedited copy of the machine output.
For the export, paste the formatted text into a word processor and use real styles: a heading style for section titles and a normal style for paragraphs. That keeps the structure when exporting to PDF, makes the document navigable for screen reader users, and lets you regenerate a table of contents. For publishing on a website, keep headings as real HTML headings and timestamps as plain text or links to the moment in the player. For research, many qualitative analysis tools import plain text or DOCX; check whether yours wants speaker labels in a specific form.
Pitfalls when formatting transcripts
- Cleaning up the words of a quote you intend to publish. Format around quotes; do not polish them.
- Mixing conventions halfway through, such as switching from names to initials, or from [00:04:10] to 4:10.
- Trusting the segment breaks as sentence breaks.
- Guessing a speaker's identity from content when the voices are similar. Mark it as uncertain instead.
- Formatting a transcript that is not yet proofread, then correcting text inside a carefully laid-out document.
- Leaving personal details in a transcript that will be shared or published.
Working from mydubly transcripts
A mydubly transcript job gives you the plain-text transcript and a timestamped transcript with one recognized segment per line and its start time in brackets, as described above, plus SRT and VTT subtitle files. It costs 1 credit per minute with a 5-credit minimum per file, so a 45-minute interview costs 45 credits (4.5¢). Whisper-based recognition tends toward clean output with many fillers dropped, which suits published and journalistic transcripts and is a limitation for strict verbatim work.
What mydubly does not do: it adds no speaker labels, headings or bracketed notes, and it does not export to DOCX or PDF. Those steps happen in your editor. You can also request a translated transcript in one of 21 languages; format it with the same speaker labels and timestamps as the original so the two can be read side by side. Start from video to text for video files.
Next step
Take one recording you need to share, transcribe it, and format the first ten minutes using the steps above. Write the conventions you chose on a single page; that page becomes your house style, and the next transcript goes twice as fast. For interviews specifically, the interview transcription page covers the recording side.
Frequently asked questions
Should I correct grammar when formatting a transcript?
It depends on the style. A clean transcript may fix false starts and obvious slips while keeping the speaker's words and meaning. A verbatim transcript keeps everything as spoken. Formatting itself, meaning paragraphs, labels and timestamps, never requires changing words, so decide the style separately and note it in the header.
How should I format a transcript for a thesis appendix?
Follow your department's guidance first. A common approach is anonymized speaker codes, a timestamp at every turn, and line numbers so you can cite passages precisely in the text, such as P3, lines 112 to 118. Keep the code key separate from the appendix, and state the transcription style and method in your methods chapter.
Can formatting be automated?
Parts of it. A spreadsheet or a short script can thin timestamps, join segments into paragraphs and apply a header template. Speaker labels and notes on unclear audio need a person listening, because the transcript does not record who spoke or how confident the recognition was. Budget human time for those steps.
Where should timestamps go in a podcast transcript?
For a published podcast transcript, timestamps at each section heading or every few minutes are usually enough, since readers skim rather than verify. If you also publish chapters, reuse the same times so the transcript and chapter list agree. Turning a transcript into show notes is covered in podcast show notes from a transcript.
How do I format a transcript with its translation?
Use a two-column table, original on the left and translation on the right, one row per paragraph or speaker turn, with the same timestamp and speaker label on both sides. This layout makes it easy for a bilingual reviewer to check the translation and for readers to find the original wording of a quote.