The four ways to source localization
Most teams end up choosing among four setups. Each puts the work, cost and responsibility in a different place.
- Full-service localization agency
- Project management, translators, reviewers, voice talent, studio and QA under one contract
- Individual freelancers
- Translators, subtitlers or voice actors you hire and coordinate directly
- AI tool used in-house
- Software produces transcripts, subtitles or dubbed drafts; your team runs it and reviews the output
- Hybrid
- An AI draft followed by professional post-editing, or an agency that uses AI internally with human review
The production steps themselves, from scripts to final checks, are covered in the video localization checklist. This article is about the earlier decision: who does that work, and how to choose them. If your main question is whether machine output can match human quality at all, start with AI vs human video translation.
A decision guide by situation
No option wins everywhere. These situations point in different directions:
- Many languages at once, broadcast or advertising quality, union or professional voice casting, tight brand control: an agency carries the coordination load you would otherwise take on.
- One or two languages, steady volume, a team member who can manage people: a small group of trusted freelancers often gives you more consistency and closer contact with the actual linguists.
- Internal training, support content, research recordings, or a large back catalog where speed and cost matter more than polish: an AI tool, with spot checks or full review by a fluent colleague.
- Content that is public but not high-stakes, such as tutorials or webinars: an AI draft plus professional post-editing balances cost and quality. The article on machine translation post-editing explains what that review involves.
- Legal, medical, safety-critical or regulated content: qualified human translators and reviewers, whatever tools they use. An AI draft can be a starting point, never the final word.
Many organizations use more than one setup: an agency for the flagship campaign and an AI tool for the help center.
Questions to ask any vendor
Whether you are talking to an agency or a freelancer, the same questions reveal how they work. Ask for specific answers, not brochures.
- Who will translate and review our content, and are they native speakers of the target language? Are they employees or subcontractors?
- What subject areas do the linguists specialize in? Can we see samples in our field?
- Do you use machine translation or AI voices anywhere in the process? Where, and with what human review?
- How do you handle terminology and style? Will you build and maintain a glossary and style guide for us?
- What does your quality process look like, and what happens when we find an error after delivery?
- How do you handle on-screen text, graphics and subtitles burned into the picture?
- Who owns the deliverables, glossaries and translation memories at the end of the contract?
- How do you price revisions, rush jobs and small updates to already-localized videos?
For AI tools, ask the equivalent questions of the product: which languages it supports, what it outputs, what it cannot do, where data goes and how long it is kept.
Deliverables to put in writing
Disagreements with localization suppliers often come down to unstated deliverables. Specify exactly what you will receive for each video and language:
- Subtitle files and their format, such as SRT or VTT, plus any line-length or reading-speed rules.
- Dubbed or voiced video files, and the dialogue and music-and-effects stems separately, so you can re-dub the dialogue and remix later.
- Translated transcripts or scripts, so you can reuse text in articles and help content.
- Glossaries, style guides and translation memory exports, so you are not locked in.
- Graphics and on-screen text versions, if in scope.
- A review or QA report listing what was checked.
Keeping translated scripts and a translation memory has long-term value; the article on translation memory explains why it matters when you update videos later.
How to judge a quality process
Every supplier says it delivers quality. Ask how. A credible process names the steps and who performs them: translation by one linguist, review by a second, in-context review of subtitles or audio against the picture, and a final technical check of timing, sync and file specs. Ask how many revision rounds are included and how feedback from your in-country reviewers is incorporated.
For AI tools, quality depends on your process rather than the vendor's. You need a reviewer who is fluent in each language, and the article on finding a reviewer for AI translation covers where to find one.
Confidentiality and data handling
Unreleased products, internal training and customer recordings are sensitive. Ask every supplier:
- Will you sign our NDA, and does it bind your subcontractors?
- Where are files stored, who can access them and how long are they kept after the project?
- Do you or your tools use our content to train AI models?
- Can we have files deleted on request, and will you confirm it?
For an AI tool, read its privacy policy and retention terms yourself rather than relying on marketing pages. If your organization has formal security requirements, involve your security team early.
Pricing models you will meet
Pricing structures vary widely, so compare like with like. You will commonly see per-minute pricing for subtitles and dubbing, per-word pricing for translating scripts, per-project quotes for complex work, separate fees for voice talent and studio time, minimum charges per order and surcharges for rush delivery or rare language pairs. Ask what is included: review, revisions, file preparation and project management can be priced separately.
AI tools usually charge per minute of media or by subscription. The video translation cost guide breaks down what drives cost across approaches. Whatever the model, estimate the total for a realistic year of work, including updates, not only the first project.
A company selling compliance courses needs its 40 course videos, about eight minutes each, in Spanish and French. It gets a quote from an agency covering translation, two-step review and professional voice-over, and separately tests an AI tool on two videos. The training lead finds the AI drafts good enough for the course's supplementary videos once a fluent contractor post-edits the subtitles, but wants professional voice-over for the six flagship lessons sold to large clients. The final setup is hybrid: the agency voices the six flagship videos, and the remaining 34 get AI-generated subtitles and dubbed drafts reviewed by a freelance translator in each language.
Running the selection
- List your content by risk and visibility: what is public, regulated or brand-critical, and what is internal or supplementary.
- Estimate volume per language for the next year, including updates.
- Decide which setups fit each content group using the guide above.
- Shortlist two or three candidates per setup and send them the same brief and the questions above.
- Run a paid pilot on the same short, representative video, ideally one with terminology and some on-screen text.
- Have your in-country reviewers score each pilot blind on accuracy, terminology, tone and timing.
- Compare total cost, turnaround and how each candidate handled feedback, then contract with clear deliverables.
When an AI tool is enough, and when it can't be
An AI tool is often enough when the content is internal or supplementary, when speed matters more than polish, when a single narrator speaks clearly and when someone fluent can review. It is usually not enough when the content is regulated or safety-critical, when several speakers need distinct voices, when lip-sync matters, when on-screen text carries meaning, or when a language you need is not supported. In those cases a human team, or a hybrid with substantial human work, is the realistic choice.
Where mydubly sits among the options
mydubly is an AI tool in the third row of the table above. It turns a video or audio file uploaded in the browser into a transcript, SRT and VTT subtitles, a translated transcript, or a dubbed version with a translated voice track and video, in 21 languages with eight stock voices. It does not provide project management, human translators, reviewers or voice actors, and it has no voice cloning, lip-sync, per-speaker voices or on-screen text translation. See video localization for how it handles the localization workflow.
On data handling, the video file stays on your device and only compressed audio chunks are sent over HTTPS. Results are deleted within 30 minutes of completion, and the privacy policy says user data is not used for training; details are on the private video translation page. Pricing is per minute: in the example above, subtitles for one eight-minute video cost 8 credits (0.8¢) and dubbing it costs 400 credits (40¢) per language.
Making the call
Start by sorting your videos into high-risk and everyday content; that split usually makes the choice obvious for each group. Then run one paid pilot per candidate on the same video. If you want to see what an AI tool produces for your own content, the video localization page explains the outputs.
Frequently asked questions
Is an agency always more accurate than an AI tool?
Not automatically. An agency's accuracy depends on the linguists and reviewers it assigns, and some agencies use machine translation internally. An AI tool's output depends on audio quality, subject matter and language, and is usually checked by your own reviewer. For regulated or high-stakes content, insist on qualified human translation and review whoever you hire, and test any supplier with a pilot before deciding.
Should we let vendors use machine translation on our content?
It can be reasonable if the vendor tells you where it is used and a qualified linguist post-edits the output, and if your confidentiality terms allow content to be processed by the tools involved. Put the policy in the contract: whether machine translation is allowed, for which content, with what review, and whether any tool may retain or train on your files.
What should a localization pilot include?
Use a short, representative video of a few minutes with real terminology, a bit of on-screen text and the kind of speaker you usually have. Give every candidate the same brief and glossary, and pay for the work so you get their normal process. Ask fluent reviewers to score the results without knowing which supplier produced which version, then compare quality, turnaround and how each handled your corrections.
Who should own glossaries and translation memories?
You should, ideally. They are built from your content and make future work faster and more consistent, whoever does it. Write into the contract that glossaries, style guides and translation memory exports are delivered to you in standard formats at the end of each project or on request. Without that, switching suppliers later can mean rebuilding terminology from scratch.
How do we compare per-word and per-minute quotes?
Convert both to a cost per finished video minute for a realistic sample. A minute of fast, dense narration has many more words than a minute of slow demonstration, so per-word pricing varies by content. Add voice talent, studio, project management, revisions and minimum charges, which are often quoted separately. Then compare the yearly total including expected updates, not just the price of the first video.
Can we mix suppliers for different languages?
Yes, and many teams do, for example an agency for languages where they lack contacts and freelancers where they already have trusted linguists. The risk is inconsistency, so share one glossary and style guide with everyone, keep the source scripts in one place and use the same review criteria for each language. Make one person responsible for coordinating across suppliers.