How text-based editors link words to time
Behind a text-based editor sits a transcript where each word has timing information: when it starts and when it ends in the recording. Speech recognition can produce these word timings directly, and some tools refine them with forced alignment, which matches a known text against the audio to place each word more precisely.
Once every word has a time range, the transcript becomes a control surface for the media. Selecting text selects the matching time range. Deleting text creates a cut that removes that range. Moving a block of text reorders the underlying clips. The editor keeps a list of these ranges, essentially an edit decision list built from your text, and renders the result as a normal sequence.
Good editors do some invisible work at each cut: they snap to gaps between words, add a few frames of handle so consonants are not clipped, and may apply short audio crossfades to avoid clicks. That is why the same text edit can sound clean in one tool and choppy in another.
What it is good for
Text-based editing shines where the words are the content:
- Rough cuts of interviews and podcasts. Read the conversation, delete tangents and repetition, and you have a first assembly in a fraction of the time spent scrubbing a timeline.
- Removing filler and false starts. Many tools can find and remove words like um and uh, repeated words and long pauses in one pass.
- Finding soundbites. Searching the text for a phrase is faster than listening for it, and you can pull selects into a new sequence.
- Restructuring. Moving an answer earlier, or combining two takes of an intro, is a drag in text.
- Collaboration with non-editors. A producer or guest can mark what to cut in a document without opening editing software.
For long recordings, searching a transcript to locate moments is useful even without a text-based editor; finding clips in long videos covers that workflow.
What it does not do well
The method has clear edges, and knowing them saves frustration:
- Picture is an afterthought. Every text cut is also a picture cut, which means jump cuts on a single camera. You still need b-roll, zooms or a second angle to cover them, and that work happens on a timeline.
- Pacing is invisible in text. A transcript does not show how long a pause felt, how a laugh landed, or whether a cut breaks the rhythm. Rough cuts made in text almost always need a pass by ear.
- Over-cleaning sounds unnatural. Removing every filler word and breath can make speech sound breathless and robotic. Keep some pauses.
- Recognition errors spread. If a word is transcribed wrong, searching for it fails, and a cut placed by text may land a word early or late in fast or mumbled speech.
- Overlapping speech is hard. When two people talk at once, word timings blur and a clean cut may not exist.
- Visual-led content gains little. Montages, tutorials driven by on-screen action, music videos and footage without much speech have little to edit as text.
Preparing footage and transcripts for text editing
The quality of text-based editing tracks the quality of the transcript, which tracks the quality of the audio. A little preparation goes a long way:
- Record each speaker on a separate microphone where possible, and keep the dialogue clean of music during recording.
- Transcribe the full recording, not a pre-trimmed version, so timings refer to the original media.
- Correct names, product terms and numbers before you start cutting; proofreading an AI transcript gives a fast method.
- Decide whether you want a verbatim transcript, with fillers and false starts, or a clean one. For editing, verbatim is more useful because it shows exactly what you can cut; see verbatim versus clean verbatim.
- Add speaker names if your tool supports them, or mark them by hand in long multi-person recordings. Speaker diarization explains how tools attempt this automatically.
The paper edit: text editing without a text editor
Long before software linked words to frames, documentary editors made paper edits: printing transcripts, marking the lines they wanted with timecodes, and handing the plan to the editor. The method still works with any editing software and any timestamped transcript.
Read the timestamped transcript, highlight the passages you want and note the times beside them. Arrange the highlights in a document in the order you want the story to run. Then build the cut on the timeline, jumping straight to each marked time. Segment-level timestamps are enough to locate each passage; the precise in and out points are set by ear in the editor.
Many editing programs can also import an SRT subtitle file as a caption track. Placed on the timeline, it shows the spoken text above the clips at the right times, which makes it easy to find lines visually while cutting in a conventional editor.
Example: a 50-minute interview to a 12-minute episode
A podcaster records a 50-minute video interview with two microphones. They transcribe it, spend twenty minutes correcting the guest's company and product names, then read the transcript and highlight eleven passages with their timestamps. In the document, they reorder the passages so the origin story comes first. In the editor, they jump to each marked time, cut the passages out with a little room either side, and assemble a 14-minute rough cut. A listening pass trims long pauses and two weak answers, and b-roll covers the jump cuts. The transcript for the full 50 minutes costs 50 credits (5¢).
Common mistakes and trade-offs
- Trusting the text cut without listening. Always play every edit point; a cut that reads perfectly can clip a breath or a word ending.
- Cutting before correcting the transcript. You search for a name and miss three mentions because they were transcribed wrong.
- Removing all pauses. Speech needs some air; leave natural breaths and short pauses, especially before important points.
- Forgetting the picture. Plan cover shots for every cut in a single-camera interview, or accept visible jump cuts as a style.
- Editing the transcript of an already-edited video. If you plan to cut from the transcript, transcribe the raw recording so timings match your source media.
The trade-off is speed against control. Text-based editing gets you to a structured rough cut quickly; a timeline gives you the fine control over rhythm, picture and sound that a finished video needs. Most editors who use text tools use both.
What mydubly provides, and what it does not
mydubly is not a video editor. It does not cut video, link text to frames, remove filler words or render an edited sequence. What it provides is the input to a paper edit or a caption-based workflow: a plain transcript, a timestamped transcript with [m:ss] labels, and SRT and VTT files in which each recognized segment has a start and end time. These timings are per segment rather than per word, so they locate passages rather than set exact cut points.
You can transcribe video or audio files up to 2 hours each in the video to text tool, in any of 21 languages, with the spoken language detected automatically. Speech recognition uses a voice activity filter that skips silence while keeping timestamps tied to the original audio, so the times still line up with your source file. There are no speaker labels, so add speaker names yourself in multi-person recordings. A transcript costs 1 credit per minute with a minimum of 5 credits; a 30-minute recording is 30 credits (3¢).
Next step
Take one interview or podcast recording, get a timestamped transcript, and make a paper edit before you open the timeline. If it saves you time, consider a dedicated text-based editor for future projects; if your content is mostly visual, a timeline and a searchable transcript may be all you need.
Frequently asked questions
Does text-based editing work for any video?
It works best when speech carries the content, as in interviews, podcasts, talking-head videos and lectures. For montages, action-driven tutorials or music-led videos, there is little text to edit and a timeline is more efficient.
How accurate are cuts made by deleting words?
They depend on how precisely the tool places word boundaries. In clear, steady speech cuts are usually clean; in fast or overlapping speech a cut can clip a syllable or leave a fragment. Listen to every edit point and adjust by a few frames where needed.
Should I remove all filler words automatically?
Usually not. Removing every um, uh and pause can make speech sound rushed and unnatural. Remove the distracting ones, keep natural breaths and short pauses, and listen to the result before applying it to a whole episode.
What is a paper edit?
A paper edit is a plan for a cut written on a transcript: you highlight the passages you want, note their timestamps and arrange them in order. The editor then builds the sequence from that plan. It works with any editing software and any timestamped transcript.
Can I import a subtitle file into my editing software to help find lines?
Many editing programs can import SRT files as a caption track, which shows the spoken text on the timeline at the right times. That makes it easier to find passages while editing conventionally. Check your editor's documentation for its supported caption formats.
Does mydubly give word-level timestamps?
Its outputs are timed by segment: the timestamped transcript uses [m:ss] labels and the SRT and VTT files have one timed cue per recognized segment. That is suitable for locating passages and building a paper edit. Precise word-by-word cut points need a dedicated text-based editor.