About this collection
Articles are grouped by topic. Each one explains how the technology works, what affects quality, its honest limitations and, where relevant, how mydubly handles it.
How the blog is organised
The articles are grouped into topics, from how the technology works to the jobs people use it for. Roughly, the topics fall into four bands. The first explains the technology: AI video translation, speech recognition, machine translation, AI voices, and how mydubly itself is built. The second is about quality: translation quality in depth, difficult recordings, audio engineering for speech, video engineering and media formats. The third is about doing the work: preparing files, workflows, repurposing recordings, troubleshooting, subtitles and captions in depth, and accessibility. The fourth covers who uses it and why: creators, education and research, business video, language learning, localization and international content strategy.
Each article owns one question and links to its neighbours rather than repeating them, so following links from any article is a reasonable way to read a topic.
Where to start
- New to AI video translation: what AI video translation is, then what it can't do yet.
- Reviewing machine output: proofreading an AI transcript, back translation for a language you can't read, and the AI dub quality checklist.
- Getting a better recording: the five-minute audio check and background noise and speech recognition.
- Fixing a file that misbehaves: unsupported video files, audio drifting from the picture and SRT files that won't load.
- Planning across languages: localization strategy and choosing languages.
- Privacy questions: what a cloud speech pipeline receives.
How articles treat mydubly
Most articles explain a subject in general terms first, the way a textbook chapter would, and only then say where mydubly fits. When the tool doesn't do something an article discusses, such as cloning a voice, matching lip movements, labelling speakers or translating text in the picture, the article says so plainly and describes what people do instead. Prices are always quoted from the published per-minute rates, and articles avoid invented statistics: where a number can't be sourced, you'll find a method for measuring it on your own files instead.
New articles are added to an existing topic rather than as standalone posts, so a topic's section on this page stays the best entry point for that subject.
Articles, guides and tool pages
Articles explain why things work the way they do and what affects the result. When you're ready to do a specific task, the guides walk through it step by step, and the tool pages, such as the video translator or audio to text, are where the work actually happens.
AI video translation
What AI video translation is, how the pipeline works, and what affects speed and accuracy.
- AI video translation, explained: what you get and who it's for
- Inside an AI video translation pipeline, stage by stage
- Should a machine or a professional translate your video?
- Judging the accuracy of AI video translation on your own footage
- Video translation turnaround: what takes time and how to plan for it
- Translating hour-long and two-hour videos with AI
- Translating interviews, panels and podcasts with more than one speaker
- Spanish for Mexico or Madrid? Adapting translated video to a regional market
- Background Music in Translated Videos: What Stays and How to Perfect It
- What AI Video Translation Actually Gets You, and When It Doesn't
- What AI Video Translation Can't Do Yet, and What to Do Instead
- Producing Videos That Translate Cleanly Later
- Slides, Lower Thirds and UI: Translating the Text in Your Picture
- Code-Switching, Bilingual Speakers and Videos That Change Language
- The A-to-Z of Video Translation Terms
Speech recognition & transcription
How speech-to-text models such as Whisper turn audio into timestamped text, and where they struggle.
- Automatic speech recognition, explained from Audrey to Whisper
- How speech to text works, from waveform to punctuated sentence
- OpenAI Whisper: how the open speech recognition model works
- Whisper large-v3-turbo: a faster Whisper with a four-layer decoder
- End-to-end Whisper versus hybrid HMM/DNN speech recognition
- Word error rate explained, with a worked example
- Transcribing long audio files: chunking, boundaries and the 2-hour limit
- How background noise, echo and music affect speech recognition
- Why Speech Recognition Struggles With Some Accents, and What Helps
- How One Speech Model Transcribes Many Languages, and Why Quality Varies
- How Speech Models Work Out Which Language Is Being Spoken
- A Faster Way to Proofread AI Transcripts Without Rereading Every Line
- Speaker Diarization: How Software Works Out Who Spoke When
- When Whisper Writes Words Nobody Said: Hallucinations Explained
- True Verbatim, Clean Verbatim or Edited: Choosing a Transcription Style
AI translation
How neural machine translation works, how quality is measured, and how to fix common errors.
- AI translation, explained: how a machine turns one language into another
- Inside neural machine translation: encoders, decoders, beam search and back-translation
- Large language models or dedicated NMT engines: which should translate your content?
- Why spoken language is harder to translate than written text
- What a translation engine can't see: pronouns, gender, formality and meaning across sentences
- The machine translation errors reviewers find most, with an example of each
- Measuring translation quality: BLEU, chrF, COMET, MQM and a spot check you can actually run
- Names, brands, jargon and acronyms: getting specialist terms right in translation
- When translations grow or shrink: what text expansion does to subtitles and dubbed audio
- Fixing machine translation efficiently: light and full post-editing for transcripts and subtitles
AI voice & dubbing
Text-to-speech, voice generation and dubbing: how synthetic voices are made and what makes them sound natural.
- Text to Speech, Explained: From Robot Voices to Neural TTS
- Inside a Synthetic Voice: The Neural TTS Pipeline Step by Step
- Chatterbox TTS: Resemble AI's Open-Source Voice Model, Explained
- How AI Voice Translation Turns Speech into Another Language
- AI or Human Dubbing? How to Decide for Each Project
- Dubbing, Voice-Over or Narration? The Differences That Matter
- AI Lip Sync Explained: Re-Rendering Mouths to Match New Audio
- Picking the Right AI Voice: Content, Audience, Pace and Testing
- Cloned Voices or Preset Voices? How to Decide for Dubbed Video
- What Separates a Convincing AI Voice from a Robotic One
- Why AI Dubbing Is Harder Than It Looks
- Using Synthetic Voices Without Deceiving Anyone
- Why Synthetic Voices Sometimes Stress the Wrong Word
- How One TTS Model Speaks Many Languages
- Quality-Checking an AI Dub Before It Goes Live
Video localization
Planning, adapting and measuring video for audiences in other languages and cultures.
- Video localization, explained: what it covers and who needs it
- Translation or localization? How to tell what your video needs
- Building a video localization strategy that scales past the first video
- How to pick the languages your videos should be translated into
- Adapting a video for another culture: humor, idioms, visuals and formats
- Localizing product demos: translated narration over an untranslated interface
- Translating town halls, leadership messages and policy videos for every office
- Localizing explainer and classroom videos: terms, reading level and captions
- Localizing video into Arabic and Hebrew: getting right-to-left text right
- Is localization working? Metrics, comparisons and when to expand or stop
Creators, YouTube & podcasts
Workflows for creators who want one recording to reach audiences in several languages.
- How YouTube's Multi-Language Audio Tracks Work, and How to Make One
- One Channel or Several? Structuring a Multilingual YouTube Presence
- Translating Your YouTube Titles and Descriptions Alongside Subtitles
- Publishing a Podcast in Several Languages: Feeds, Naming and Cadence
- One Recording, Many Languages: Planning a Repurposing Workflow
- From Transcript to Article: Turning a Video into a Blog Post
- Captions and YouTube Discoverability: What Is Known and What Is Guesswork
- Translating Your YouTube Back Catalog: What to Do First
- Translating Stream VODs: Long, Noisy and Full of Chat
- Writing Show Notes, Chapters and Summaries from an Episode Transcript
- Translating Gaming Videos: Commentary, Game Audio and Gamer Slang
- Community Subtitles After YouTube's Feature: Organizing Volunteer Translators
- Translating Someone Else's Video: Copyright Basics Before You Publish
- Thumbnails in Other Languages: Text, Layout and Cultural Cues
- Planning a Multilingual Content Calendar: Release Timing, Review and Tracking
- Reading, Replying to and Moderating Comments in Languages You Don't Speak
- What to Do With Sponsor Reads When You Translate or Dub a Video
- How to Translate a Video Essay Without Losing Its Argument
- Translating Fiction Podcasts and Audio Dramas Into Another Language
Education & research
Transcripts, captions and translation for teaching, studying and research.
- How Students Can Use AI Transcription Without Cutting Corners
- Make a Lecture Archive Searchable with Transcripts
- Dubbed Audio, Subtitles or Transcripts? Videos for Students with Other Home Languages
- Captions and Transcripts That Make Course Videos Accessible
- A Practical Method for Learning a Language from Video
- Transcribing Oral History Interviews for the Archive
- From Recorded Lecture to Study Notes You Will Actually Use
- Reaching Every Family: Translating School Videos into Home Languages
- Getting Usable Transcripts from Focus Group Recordings
- Quoting and Citing Video and Audio Sources with Timestamps
Privacy & architecture
Where your media goes during processing, and why keeping video on the device matters.
- How Browsers Learned to Process Video Without a Server
- What On-Device Processing Protects in a Video, and What It Doesn't
- Processing Video in the Browser or on a Server: The Real Trade-offs
- What Happens to Your Audio in a Cloud Speech Pipeline
- From Gigabytes to Megabytes: Shrinking What a Video Translation Uploads
Comparisons
Side-by-side explanations of terms and approaches that are easy to confuse.
- Transcription or Translation? How They Differ and Fit Together
- Captions and Subtitles Are Not the Same Thing
- Burned-In or Selectable? Choosing Between Open and Closed Captions
- Machine or Person? Comparing AI and Human Transcription
- One Voice or Many? Choosing a Dubbing Approach
- MP3 or WAV for Transcription? It Depends on What You Do Next
- MP4 vs MOV: Two Video Containers and When Each Makes Sense
- Transcript or Subtitles: Which Text Does Your Video Need?
- Audio Description and Subtitles Serve Different Viewers
- AAC or MP3? How Two Lossy Audio Codecs Differ in Practice
- MKV vs MP4: What Each Container Can Hold and Where It Plays
- H.264 vs H.265: Choosing Between AVC and HEVC for Each Destination
- Dictation or Transcription? Speaking Text Live vs Converting a Recording
- Internationalization vs Localization: Preparing Content, Then Adapting It
- Translation, Transliteration and Transcription: Meaning, Letters and Sounds
- Embedded Subtitles, Sidecar SRT Files or Burned-In Text: Where Each Works
- Speech Recognition vs Voice Recognition: What Was Said, Who Said It, and When
How it's built
The engineering behind long-file processing, timing alignment and in-browser video assembly.
- How to Split Long Audio for Speech Recognition Without Cutting Words in Half
- From Recognition Segments to Subtitle Cues That Stay in Sync
- Keeping a Translated Voice in Time With the Original Speaker
- Replacing a Video's Audio Track in the Browser Without Touching the Picture
- Working With Multi-Gigabyte Video Inside a Browser Tab
Translation quality in depth
The hard cases for translated video: idioms, register, gender, numbers, lyrics, and how to check a translation you can't read.
- How to Write a Translation Style Guide for Your Videos
- Pivot Language Translation: Why So Much Translation Passes Through English
- Translating Children's Videos Into Other Languages
- How to Translate Swearing and Sensitive Language in Subtitles
- Translating Song Lyrics in Videos: Subtitles, Singable Versions and Rights
- Back Translation: Spot-Checking a Video Translation You Can't Read
- Finding a Reviewer for an AI Translation, and Briefing Them Well
- Translation Memory Explained: How Localization Teams Reuse Approved Translations
Difficult recordings
Fast, quiet, clipped, windy, phone-line and archival audio: what each problem does to transcription and how to deal with it.
- Recording With Several Microphones: How to Mix for Transcription
- Transcribing Fast Talkers, Commentators and Sped-Up Recordings
- Getting a Usable Transcript From Quiet or Distant Audio
- Transcribing Recorded Phone Calls: Narrowband Audio, Consent and Exports
- From Cassette to Transcript: Digitizing Old Tapes for Transcription
- Clipped Audio: What It Does to Speech and How Far It Can Be Repaired
- Wind Noise in Outdoor Video: Why It Happens and What You Can Still Save
- Transcribing Recordings of Conferences, Ceremonies and Public Meetings
Preparing your files
Recording, extracting, splitting, converting and checking media before you transcribe or translate it.
- How to Extract Audio from a Video, and When You Don't Need To
- How to Split a Video into Parts Without Losing Quality
- Converting MOV to MP4 and Other Video Formats Without Wasting Quality
- Screen Recording with Audio: Microphone, System Sound and Clean Narration
- Recorder Settings for Speech: What to Set and What to Leave Alone
- How to Record Audio on Your Phone So It Transcribes Well
- Recording a Remote Interview over a Video Call, Cleanly and with a Backup
- A Five-Minute Audio Quality Check Before You Transcribe
- Choosing a Microphone for Speech and Transcription
Workflows
Step-by-step ways to get from a recording to the finished deliverable, from editing SRT files to multi-language releases.
- How to Build One Video File with Original and Translated Audio
- How to Show Subtitles in Two Languages at Once
- Editing SRT Subtitle Files by Hand and in Subtitle Editors
- Finishing an AI Dub in Premiere, Resolve, Final Cut or a Free Editor
- A Review Workflow for Teams Translating Videos
- Formatting a Raw Transcript So People Can Actually Read It
- Working With Footage in a Language You Don't Speak: A Reporter's Workflow
- When You Need a Certified Translation of a Recording, and Where AI Fits
- Running One Video Through Several Languages Without Losing Track
- Translating an Audiobook or Long Narration, Chapter by Chapter
Repurposing recordings
Turning transcripts of videos, meetings and interviews into articles, newsletters, chapters, clips and help content.
- How to Turn a Video into an Email Newsletter
- From Customer Interview to Written Case Study
- Turning Product Walkthrough Videos into Help Center Articles
- Turning a Video Transcript into Text Posts for Social Media
- How to Publish Transcripts Alongside Your Videos and Podcasts
- Finding Short Clips in a Long Video Using Its Transcript
- Turning a Meeting Transcript into Useful Notes
- Editing an Interview Transcript into a Q&A or Feature Article
- Turning Course Videos into Written Lessons and Handouts
Language learning
Methods for learning languages with transcripts, podcasts, dictation and listening practice.
- Comprehensible Input with Video: The Theory, Its Critics and a Practical Method
- Turning Video Transcripts into Vocabulary Lists and Audio Flashcards
- How to Learn a Language with Podcasts, from Learner Shows to Native Ones
- Dictation Practice: Write Down What You Hear and Learn from Every Mistake
- Following University Lectures in a Language That Isn't Your First
- Can Text to Speech Be Your Pronunciation Model? Uses, Limits and Shadowing
- Using Authentic Video in ESL and EFL Lessons, from Clip Selection to Copyright
- Extensive Listening: How to Build Hours of Easy Listening into Your Week
- Using AI Translation in Language Classes Without Short-Circuiting Learning
Business video
Translating support, sales, safety, healthcare and research videos, and choosing who does the work.
- Running a Support Video Library in More Than One Language
- Translating Safety Training Videos Without Losing the Warnings
- Selling in the Buyer's Language: Translating Sales Videos
- Agency, Freelancers or AI Tool? Choosing Who Localizes Your Videos
- Product Videos for International Stores: A Translation Playbook
- Translating Patient Education Videos Safely
- Public Information Videos in Every Language Your Community Speaks
- Working With User Research Sessions Recorded in Other Languages
Audio & video formats
Codecs, containers, sample rates, bitrates, channels and loudness, explained for people working with speech.
- Audio Codecs Explained: How Sound Is Compressed and Played Back
- How Video Codecs Shrink Moving Pictures, and Why It Matters for Your Files
- Variable Frame Rate Video: Why It Happens and When to Convert It
- Audio Sample Rates From 8 kHz to 48 kHz, and Which Ones Matter for Speech
- How Much Bitrate Does Speech Need Before Transcripts Suffer?
- Mono or Stereo for Voice Recordings: What Changes When Channels Are Combined
- Loudness Normalization: LUFS, True Peak and Consistent Speech Levels
- When a Video File Has More Than One Audio Track
- The Opus Codec: One Format for Voice Calls, Music and Speech Uploads
- Why HEVC Phone Videos Don't Always Play, and How to Fix It
Troubleshooting
Common problems with sync, subtitles, transcripts, AI voices and media files: causes, diagnosis and fixes.
- Why Your Audio Drifts Away From the Picture, and How to Fix It
- What to Do When a Video Is Too Big to Send or Upload
- When an SRT Subtitle File Won't Load: A Checklist That Finds the Cause
- Why a Transcript Skips Words, Sentences or Whole Minutes
- When the Synthetic Voice Says a Name Wrong: Finding Where It Broke
- Muffled Speech in a Recording: Causes, Checks and Realistic Fixes
- Hum and Buzz in a Recording: Tracking Down Mains Noise and Cleaning It Up
- Silent After Export: Tracking Down a Video's Missing Audio
- Garbled Subtitles: Why Accents Turn Into Symbols and Scripts Into Boxes
- Unsupported Video File Errors: What the Message Really Means
- Why AI Transcripts Get Punctuation Wrong, and How to Clean It Up
- Twenty-Five or 25? Making Numbers Consistent in AI Transcripts
- Gaps, Clicks and Stutters: Tracking Down Audio Dropouts
- Sound but No Picture: How to Fix a Video That Plays Black
- Voice Only in the Left Ear? Fixing One-Sided Audio
- Why Your Video Stutters, and Whether the File Is to Blame
- Chipmunk or Slow-Motion Voice? Fixing Audio at the Wrong Speed
- Hearing Two Voices at Once: Fixing Doubled Audio in a Video
Industry & future
Where video translation, dubbing and accessibility technology came from and where it is heading.
- Where AI Video Translation Is Heading, and What Is Still Guesswork
- What AI Changes, and Doesn't, for Accessible Video
- How Real-Time Speech Translation Works, and Where Files Do Better
- A Short History of Dubbing and Subtitles, From Title Cards to AI
- AI and Sign Language: What Machines Can and Can't Translate Yet
- Speech AI on Your Own Device: What Runs Locally and What Still Needs a Server
Accessibility & inclusive video
Making videos usable for viewers who are deaf, hard of hearing, blind, low vision or neurodivergent, from captions to players and description.
- A Pre-Publication Checklist for Accessible Video
- What Makes a Video Player Accessible, and How to Test One
- Captioning Sound Effects, Music and Speakers: Conventions That Work
- Running an Accessible Webinar: Before, During and After
- Building an Accessible Online Course, Module by Module
- Making Workplace Training Videos Work for Every Employee
- How to Improve Dialogue Clarity in Video So Everyone Can Follow the Speech
- Multilingual Accessibility: Keeping Every Language Version of a Video Usable
- Writing a Plain Language Video Script That More Viewers Can Follow
- Accessible Social Media Videos: Captions, Placement, Descriptions and Safety
- Descriptive Transcripts: Putting the Whole Video into Text
- Writing an Audio Description Script, Step by Step
Subtitles & captions in depth
Reading speed, timing rules, segmentation, positioning, formats and quality control for subtitle and caption files.
- Subtitling Your Film for Festival Submission and Screening
- Professional Subtitle Spotting: Durations, Gaps and Shot Changes
- Splitting Subtitle Text Into Cues and Lines That Read Naturally
- Italics, Dashes, Ellipses and Numbers: Subtitle Typography Explained
- Where Subtitles Belong on Screen and When to Move Them
- Making Subtitles Easy to Read on Any Screen and Any Footage
- A Field Guide to Subtitle and Caption Formats Beyond SRT and VTT
- CEA-608 and CEA-708 Compared: How Broadcast Captions Travel With the Picture
- Forced Subtitles Explained: Translating Only the Lines Viewers Can't Follow
- How to Run a Quality Check on a Finished Subtitle File
- Subtitle Templates: One Timed Master File for Every Language
- Live Captioning Explained: Stenography, Respeaking and Automatic Captions
Audio engineering for speech
Bit depth, dynamics, room acoustics, EQ and microphone technique for speech that will be transcribed, translated or dubbed.
- Audio Bit Depth: What 16, 24 and 32-Bit Float Really Change
- Compressing Voice Audio: Settings, Limiters and When to Leave It Alone
- Treating a Room for Voice: Reflections, Echo and Where Absorption Goes
- How Acoustic Echo Cancellation Works, and What It Does to Recorded Calls
- Voice Activity Detection: How Software Decides When Someone Is Speaking
- What Makes Speech Intelligible, and Why Loud Is Not the Same as Clear
- How to EQ a Voice Recording: Frequencies, Cuts and Sensible Boosts
- Plosive Pops in Voice Recordings: Prevention and Repair
- Harsh S Sounds in Voice Recordings: Causes, De-Essers and Technique
- Clothing Rustle and Cable Thumps: Getting Clean Sound From a Lavalier
- dBFS and Audio Levels Explained: Meters, Headroom and Recording Targets
- Speaking for the Transcript: Delivery Habits That Make Recordings Easy to Transcribe
Video engineering & media pipelines
Demuxing, timestamps, keyframes, frame rates, bitrates and decoding: how video files and processing pipelines work.
- Demuxing Explained: How Players and Tools Pull Streams Out of a Video File
- PTS and DTS: How Video Files Record When Each Frame Is Decoded and Shown
- How Players Keep Sound and Picture Together: Master Clocks, Timestamps and Frame Drops
- Keyframes and GOP Structure: How Frame Types Shape Seeking, Cutting and File Size
- Inside an MP4 File: Boxes, Sample Tables, Fast Start and Fragments
- Video Frame Rates in Practice: 24, 25, 30, 60 and the Odd 29.97 Family
- The Media Processing Pipeline: From Probe to Package, Stage by Stage
- Video Bitrate Explained: Rate Control, File Size and Export Settings
- Hardware Video Decoding: Why Some Files Play Smoothly and Others Stutter
- Interlaced or Progressive: Fields, Frames, Combing and Deinterlacing
Knowledge work with recordings
Organizing, searching, summarizing and checking information captured in meetings, interviews and video archives.
- How to Make Months of Meeting Recordings Searchable
- How to Organize a Growing Collection of Audio Files
- Video Metadata and Tagging for a Library You Can Search
- Searching AI Transcripts When the Words May Be Wrong
- Using Transcripts in a Personal Knowledge Base Without Drowning in Them
- How to Review a Long Recording Without Listening to All of It
- Knowledge Capture Interviews: Recording What Experts Know Before They Leave
- How to Summarize a Transcript Accurately
- How to Fact-Check Claims Made in Videos and Podcasts
- How to Organize Family Videos So You Can Find Every Story
International content strategy
Governance, team structure, journeys, search and website delivery for content published in several languages.
- Planning What Every Language Gets Across Web, Docs, Email, Social and Video
- Who Owns, Approves and Retires Content in Every Language
- When the Source Video Changes: Keeping Every Translation in Step
- Making Translated Videos Findable on Your Own Website
- How to Organize the People Behind Your Localization Work
- Keeping Your Brand Voice Recognizable in Every Language
- Localizing Each Stage of the Customer Journey in Order
- Writing a Localization Project Plan From Scope to Close-Out
- Central Translation or Local Production: Choosing Per Content Type
- Offering Several Language Versions of a Video on Your Site
Research & academic workflows
Transcription conventions, qualitative analysis, fieldwork, data management and ethics for research recordings.
- From Interview Recordings to Themes: Preparing Transcripts for Analysis
- Choosing and Writing Transcription Conventions for Research Data
- Interviews in Another Language: Translation Decisions in Qualitative Research
- Recording Interviews in the Field Without Losing Data
- Managing Research Recordings and Transcripts From Capture to Archive
- Returning Transcripts to Participants: A Practical Guide to Member Checking
- Sending Research Audio to an AI Transcription Service: The Ethics Questions
- How to Build a Corpus of Spoken Language
- Doing Content Analysis on Interview, Media and Meeting Transcripts
- How Long It Really Takes to Transcribe an Interview
Creator content operations
Assets, procedures, stems, licensing and hand-offs that keep a multilingual channel running.
- Video Asset Management for Creators: Folders, Names, Masters and Backups
- A One-Page Localization SOP for Every Upload
- Publishing Language Versions of a Video Across Platforms and Your Own Site
- The M&E Track: Why Every Project Should Export Music and Effects Without Dialogue
- Does Your Music License Cover the Translated Version of a Video?
- Text-Based Video Editing: Cutting a Video by Editing Its Transcript
- Textless Masters and Elements: Exporting a Clean Picture for Localization
- Handing a Localization Edit to a Freelance Editor: The Complete Checklist