Why interviews are their own subtitling job
A scripted video has one voice and clean audio. An interview has at least two people on different microphones, answers that begin before the question ends, a subject who talks in long looping sentences, and an edit that has already removed half the conversation. Subtitles that would pass on a vlog get confusing here: the viewer can't tell whose words are on screen, a quote reads differently from how it was said, or a translated answer drifts out of step with the face delivering it.
Most of the work is editorial. AI gets the words down fast; your job is deciding what each line should say and making sure the audience always knows who is talking.
Subtitle the cut, not the raw recording
Open the finished edit, not the two-hour rushes. The editor transcribes whatever audio is in the file, so if the questions were trimmed out, there is nothing to subtitle for the interviewer, which is usually what you want. If you kept the questions but they were picked up faintly on the subject's microphone, expect the AI to miss or mangle them; type those lines yourself, or replace them with on-screen question cards in your editing app.
For a long conversation you still have to cut, get a full transcript first, choose your quotes from it, and subtitle only the locked edit. Interview transcription is built for that first step, and finding clips in long videos shows how to use the transcript to assemble the edit.
Keeping two voices apart
The AI subtitles here don't name speakers; they arrive as one stream of short lines, each beginning on the first word of a spoken phrase. Since a new phrase usually starts when someone new speaks, a change of voice tends to land on a line boundary, but not always. On your first pass:
- Watch the boundaries where the voice changes. If one line holds the end of a question and the start of an answer, put the playhead on the switch and press Split.
- Where both voices are heard but only one face is on screen, prefix the off-screen speaker's line with their name, or with a dash if that is your project's convention.
- When people talk over each other, subtitle the person the story needs and drop the crosstalk rather than cramming both voices into one unreadable line.
- Check that the subtitle changes when the face changes. A line that lingers after the cut to the other person reads as theirs.
Text colour and italics apply to the whole video, so colour-coding interviewer and guest isn't possible. A name on the first line after each change of speaker is the reliable substitute. If you wonder why software finds this hard, speaker diarization explains the problem.
Name supers and subtitles in the same frame
Documentaries and brand films usually introduce each person with a lower third, a name and title near the bottom for the first few seconds. If that graphic is already in the video, the subtitles will collide with it unless you plan for it:
- Raise the subtitles for the whole video with the vertical offset so they sit just above the lower-third zone. Simplest when the supers are short and low.
- Move the subtitles to the top. That suits interviews framed with headroom, though some viewers find it unusual.
- Render the subtitles first and add the lower thirds afterwards in your editing app, positioned around them, or deliver a subtitled master without supers.
A Box or Rounded background separates subtitles from busy graphics more cleanly than an outline alone. For restrained styling in corporate and brand work, see professional video captions.
When the interviewee speaks another language
Foreign-language interviews are where subtitles carry the whole story. Turn on AI subtitles, set "Subtitle language" to the language your audience reads, and the speech is transcribed and translated into one of 21 languages in a single pass. Translation works from what is said in the video, which matters in journalism: every subtitle is a rendering of someone's words, so a person who speaks the source language should check it before anything is quoted or broadcast.
Translated lines are split by text length, so their timing is approximate. Drag the edges so each line appears as the person begins the matching phrase, especially at moments when the viewer is watching the face. If the interviewer and guest speak different languages, review both halves closely; mixed-language audio is harder for speech recognition and errors gather where the language switches. Translating mixed-language videos and translating foreign footage for journalism go deeper, and burning translated subtitles covers the render.
How faithful should the words be?
People don't speak in finished sentences. A verbatim subtitle with every "you know", restart and trailing clause is accurate but tiring, and it often trips the warning for lines too fast to read. A heavily tidied subtitle reads smoothly but can put words in someone's mouth. A rule that works for most interviews:
- Remove fillers, stammers and repeated words that add nothing.
- Keep the speaker's vocabulary, grammar and tone, even when informal.
- Never merge two separate statements into one or reorder clauses so they mean something new.
- For a legally or journalistically sensitive quote, stay close to verbatim and accept a busier line.
Verbatim versus clean verbatim lays out the same trade-off from a transcriber's point of view.
Three interviews, three set-ups
AI subtitles in English with fillers trimmed, the Classic preset raised slightly to clear the name super. The AI pass costs the minimum of 5 credits (0.5¢) and rendering is free.
Request English subtitles, review the translation with a fluent speaker, adjust timing at the emotional beats, then export an SRT for the festival copy and render a burned-in screener.
Two people sharing a handheld mic and plenty of crosstalk. Subtitle the answers, put a name on any question kept in the cut, switch to the Social Media preset and keep text inside the 9:16 safe area; vertical video subtitles covers that layout.
Once the edit is locked and all you need is an SRT to hand to an editor, the subtitle generator is the quickest route.
About the subtitle editor on this page
- Opens
- MP4, MOV, M4V, WebM, MKV videos your browser can decode
- Subtitles from
- AI (from the speech, optionally translated into one of 21 languages), an SRT or VTT file, or typing them in
- You download
- An MP4 with the subtitles drawn into the picture at the original resolution and frame rate, plus SRT and VTT files
- Cost
- Editing, styling, rendering and SRT/VTT export are free, with no account and no watermark. AI subtitles cost 1 credit per minute (5-credit minimum per video); signing in with Google gives 100 free credits and accounts get 10 free credits a day
- Privacy
- The video never leaves your device and is rendered in your browser. Only for AI subtitles is the audio sent, over HTTPS, and it is deleted within 30 minutes of delivery
- Rendering needs
- Chrome or Edge 94+, Safari 16.4+, or Firefox 130+ on a computer. Editing and SRT/VTT export work in any current browser
More: subtitle and caption generators
- Video Caption Generator With Styled, Burned-In Captions
- Auto Subtitle Generator for Video
- AI Caption Generator: How It Works and How to Check It
- Caption Maker for Short Clips, Quotes and Memes
- Closed Caption Generator for SRT and VTT Files
- Create SDH Subtitles for Deaf and Hard-of-Hearing Viewers
- Add Captions to Lecture Videos
- Captions for Video Ads That Play Without Sound
- Add Subtitles to Gaming Videos Without Covering the HUD
Subtitle tools
- Add Subtitles to a Video
- Burn Subtitles Into a Video
- Online Subtitle Editor
- SRT to MP4: Put Your Subtitle File Into the Video
- VTT to MP4: Burn WebVTT Captions Into a Video
- Add Captions to a TikTok Video
- Add Subtitles to Instagram Reels
- Add Subtitles to a YouTube Video
- Add Hindi Subtitles to a Video
- All subtitle tools
- AI subtitle generator (SRT and VTT files)
Frequently asked questions
Does the AI tell the interviewer and the guest apart?
No. The AI subtitles contain the spoken words without speaker names, so add a name or a dash at the start of a line wherever the speaker isn't obvious from the picture.
How do I subtitle only the interviewee and not the questions?
If the questions were cut from the edit, there is nothing to subtitle. If they are still audible, delete those lines in the line list after the AI pass, or show them as on-screen question cards in your editing app.
Can I put English subtitles on a foreign-language interview?
Yes. Set the subtitle language to English before running AI subtitles and the speech is translated as it is subtitled. Have a speaker of the original language check quotes before publication, because machine translation can miss tone and idiom.
Will the subtitles clash with my lower-third name graphics?
They can if both sit at the bottom. Raise the subtitles with the vertical offset, move them to the top, or add the lower thirds after rendering in your editing software.
Can I give the subtitles to my editor instead of burning them in?
Yes. Export SRT or VTT and import it into Premiere Pro, DaVinci Resolve, Final Cut Pro or another editing app, where it can be restyled to match the rest of the film.