Why it helps
Typing up a one-hour interview takes several hours. A transcript lets you scan for the strongest quotes and check them against the recording.
Interviews are often sensitive. Audio is processed in chunks and deleted as soon as you have the results — within 30 minutes at most — and never used to train models.
Step by step
- Record with your phone or recorder close to both people — a lapel mic for the interviewee helps most.
- Drop the file (M4A, MP3, WAV) into the audio workspace with text-only output.
- Search the timestamped transcript for key phrases, then listen back to confirm exact wording.
Tips
- For phone interviews, record each side if your app allows — the remote voice is often quieter.
- Translate foreign-language interviews by choosing a target language: you get the original and an English transcript.
Example
A freelance journalist transcribes a 47-minute M4A interview for 47 credits (4.7¢) and pulls six quotes in ten minutes.
Cost: 47 credits (4.7¢).
Limitations to know
- No speaker labels — mark interviewer and interviewee yourself.
- Always confirm quotes against the audio before publishing.
Set up the recording so the transcript needs less fixing
The cheapest correction is the one you never have to make. Before the first question, ask the interviewee to say and spell their full name, job title and any organisation they are likely to mention. That puts the correct spelling near the top of the transcript, where you can copy it instead of guessing at what the recognizer heard.
- Put the recorder nearer the interviewee than yourself. Your questions can be paraphrased later; their answers are what you quote.
- For a call, record locally on each side if you can, then mix both voices into a single track before uploading. Remote interview recording covers double-ender setups and why headphones prevent echo.
- Jot down the clock time when something quotable is said. With a timestamped transcript, a scribbled "14:20 pricing story" takes you straight to the line.
- Try not to talk over the end of answers. Overlapping speech is where machine transcripts most often drop or merge words.
Telling interviewer and interviewee apart
mydubly transcripts do not label speakers, so the output is one stream of timestamped lines. In a one-to-one interview that is usually manageable: questions tend to be short and answers long, and the timestamped .txt keeps each segment on its own row. Open it in a text editor and prefix lines with Q: and A: as you skim. For panels or interviews with two guests, mark only the passages you intend to use; attributing every line of a 90-minute conversation by hand is rarely worth the time. The guide to formatting a transcript shows conventions for adding names manually.
Checking a quote against the tape
A machine transcript is a lead, not a source. Before a quote is published:
- Search the timestamped transcript for one distinctive word from the passage rather than the whole phrase, because a single misheard word breaks an exact-phrase search.
- Open the original recording in any player and jump to that timestamp, starting a few seconds early.
- Listen to the full sentence and the question that prompted it, so the quote keeps its context.
- Compare word by word. Whisper-based recognition tends to produce clean text and drops many fillers and false starts, which suits a print quote but not a verbatim record (see verbatim vs clean verbatim).
- Write the timestamp beside the quote in your draft so an editor or fact-checker can confirm it quickly.
Names, numbers and negatives deserve a second listen: "fifteen" and "fifty", "can" and "can't" are classic misses. The AI transcript proofreading guide has a fuller checklist.
Which download fits which job
- Finding and verifying quotes
- Timestamped transcript (.txt), searched in a text editor
- Writing a Q&A piece
- Plain transcript, edited into questions and answers (see interview to Q&A article)
- Subtitling a filmed interview
- SRT or VTT, attached to the video in your editor or player
- Archiving the interview
- The original recording and its timestamped transcript, stored side by side under the same name
Interviews recorded in another language
Pick a target language in transcript mode and the transcript and subtitles come back translated. The result screen shows the original and translated text side by side, which helps if you read a little of the source language. Both texts can be downloaded from the same run: Download transcript gives the translation and Download original transcript the wording as spoken. Treat every translated quote as unverified until a fluent speaker has checked it against the audio, and tell readers when a quote has been translated. Working with foreign-language footage covers verification and source protection in more depth.
Three recorded interviews of 22, 38 and 64 minutes cost 22 + 38 + 64 = 124 credits (12.4¢). A 3-minute follow-up call is billed at the 5-credit minimum, bringing the week to 129 credits (12.9¢).
Using the tool on this page
- The spoken language is detected automatically; you select the output language — the spoken language itself for an untranslated transcript, or another language for a translation. Files can be up to 2 hours long.
- Audio is sent over HTTPS in small chunks and deleted with its transcripts within 30 minutes of delivery (how files are protected).
- You can download a timestamped transcript, subtitles (SRT and VTT) and plain text. Every output, step and limit is explained on the audio to text page.
Frequently asked questions
Is my interview audio stored?
Only while it's processed. It's deleted as soon as your browser has the results, within 30 minutes at most. Your account keeps only the file name, length and credits used.
Can it tell who is speaking?
No. The transcript is continuous text with timestamps, without speaker names.
Will the transcript include ums, pauses and false starts?
Mostly not. Whisper-style recognition leans toward clean text and often leaves out fillers and repeated words. If your work needs a true verbatim record, add those features by hand while listening to the passages that matter.
Can I transcribe a video interview instead of an audio file?
Yes. MP4, MOV, WebM, MKV and M4V files are accepted alongside audio, and the video to text tool returns the same timestamped transcript plus SRT and VTT files. The SRT is handy if the interview will be published as a subtitled clip.
My interview runs longer than two hours. What should I do?
Split the recording at a natural break, such as a pause between topics, and upload each part on its own. Each part's timestamps start from zero, so note where part two begins in the full recording when you cite times. Splitting a video file explains how to cut without re-encoding.