Video Translation & Localization
One video, ten languages. How AI translates, dubs and syncs video — and the difference between literal translation and real localization.
Translation vs. localization
Translation converts words from one language to another. Localization adapts the content to the audience — idiom, units, currency, examples, regulations, humour. AI makes both cheap enough to do at scale, so one production can reach several audiences.
Most people online do not speak English as a first language. How much a translated version adds depends entirely on your subject and where your demand already is — so measure it on one video before committing to a pipeline.
You will see claims that translating doubles or triples an audience. Treat that as a possibility, not a forecast: it depends on whether there is demand for your topic in that language, whether the algorithm surfaces you there, and whether anyone in that market has heard of you. The way to find out is to translate one thing and look at the numbers, which costs very little.
Four layers, rising cost and rising risk
"Translating a video" describes four quite different jobs. Choosing deliberately between them is most of getting this right.
- Subtitles. Cheapest, fastest, lowest risk, and reversible — the original audio is still there, so a viewer can hear that a subtitle is off. Also the only layer that helps deaf and hard-of-hearing viewers.
- Dubbing with a stock voice. More immersive. The original is now gone, so an error is invisible to the viewer and to you.
- Dubbing with the speaker's cloned voice. Consistent vocal identity across languages, and now a specific person appears to have said things they did not say, in a language they may not speak.
- Lip-synced video. The most convincing and the most expensive, and it produces footage of a real person apparently speaking a language fluently.
Notice what rises alongside quality: how hard it is to notice a mistake. With subtitles, a bilingual viewer catches an error immediately and thinks "bad subtitle". With a lip-synced clone, the same error is your CEO confidently stating something false, in their own voice, with their own mouth. The higher you climb, the more you need review you can trust.
Errors compound down the chain
Every layer is built on the one below, so a mistake at the bottom does not stay small.
Transcription mishears a product name. Translation renders the misheard word fluently into Spanish. The voice engine says it confidently. Lip sync makes the mouth match. What began as one wrong word is now a natural-sounding sentence, spoken in a familiar voice, that nobody watching has any way to question.
This is why the fix belongs at the top of the chain, not the bottom. Correct the transcript before anything downstream runs. Five minutes fixing names and terms in the source text saves every language from inheriting the error, and it is the single highest-return habit in this whole workflow.
Build a do-not-translate list first
Before translating anything, write down two short lists. Most translation embarrassments come from missing these.
- Never translate: your product and company names, feature names, people's names, and anything trademarked. Translation engines will happily render a product name into its literal meaning, which is how a brand becomes a common noun in a market you have not entered yet.
- Always translate this way: your key technical terms, fixed per language. Pick the rendering once so the same concept does not appear three ways across a series — that inconsistency reads as carelessness to anyone in the field.
Most serious tools accept a glossary or a do-not-translate list; where they do not, correct the transcript by hand before translating. Either way this is a one-off job that pays out on every future video, and it takes about twenty minutes.
Automatic subtitles
The cheapest layer and the one to start with. AI transcribes, translates, and produces timed subtitles per language.
- Transcription — a model such as Whisper converts speech to text. Fix names and jargon here, before anything else runs.
- Translation — into each target language, with your glossary applied.
- Timing and readability — this is where machine output is weakest and where a little manual work shows most.
What makes subtitles readable
- Two lines maximum, and roughly forty characters a line. Longer lines force the eye off the picture.
- Reading speed. A subtitle has to stay on screen long enough to read, which is often longer than the sentence took to say. Automatic timing tends to match the speech, not the reader.
- Break at sense, not at width. Split between clauses rather than mid-phrase; auto-splitting does not know the difference.
- Keep out of the lower interface area, where platforms put their own controls and captions collide with usernames and buttons.
- Right-to-left languages — if you are translating into Hebrew or Arabic, check the rendered direction and punctuation placement rather than assuming the player handles it. Mixed text with Latin product names is where this breaks.
Burned in or as a file
A sidecar file (SRT or VTT) can be toggled, is indexable by the platform, and lets viewers choose their language — so upload one per language wherever the platform allows it. Burned-in captions always appear, which is what you want on social feeds that autoplay muted, but they cannot be switched off or corrected without re-rendering. Many people do both: burned-in for social cut-downs, sidecar files on the long-form upload.
Which languages, and how to decide
The instinct is to pick the biggest languages. The better move is to look at where you already have unexplained demand.
Your analytics will show viewers from places your content does not serve. That is existing pull, and translating for it converts something already happening. Adding a language with no signal behind it means also building an audience there from zero, which is a much larger project than a translation.
Start with one language, one video, and give it long enough to mean something. Two or three well-localized languages consistently maintained beat eight machine-translated ones that nobody in those markets would describe as good.
Dubbing & voice cloning
Replacing the soundtrack in another language. With ElevenLabs and similar you can clone the original voice, so the speaker "speaks" each language in their own voice and vocal identity stays consistent across versions.
- Multilingual dubbing — a natural-sounding voice per language.
- Tone and emotion carried across, not just words.
- Voice cloning — the same speaker everywhere.
Consent is not a formality here. Clone a voice only with the person's explicit, written permission, covering which languages, which uses and for how long — and what happens when they leave or withdraw it. A cloned voice outlives the working relationship, and "we had a chat about it" is not a position you want to be in later. This is the same paperwork discussed on the avatars page, and for the same reason.
One quality note: the clone reproduces the recording conditions as faithfully as the voice. Source samples from a quiet room with a decent microphone produce a clean clone; samples with room echo produce a clone that echoes in every language.
The problem nobody expects: length
The same sentence takes a different amount of time in different languages — some run noticeably longer than English, some shorter. Over a few minutes that drift accumulates, and it breaks things that were fine in the original.
Where it hurts: narration that referred to something on screen now arrives after it has gone; a cut timed to a word lands mid-sentence; a call to action appears while the voice is still mid-clause. Dubbing tools compensate by compressing or stretching speech, which sounds rushed or oddly slow when pushed hard.
Two habits make video translate well. Leave slack — do not time visuals tightly to individual words. And avoid deixis: "as you can see here" needs the thing to be on screen at that instant, whereas "the settings page has three options" survives a few seconds of drift. Writing the original script with translation in mind costs nothing and removes most of this.
The text inside the video
Subtitles and dubbing translate the audio. They do nothing about the words baked into the picture — titles, lower thirds, labels on a diagram, the interface in a screen recording, the closing slide with your offer.
A "fully translated" video where every graphic is still in English signals immediately that this is a translation rather than something made for the viewer, which undercuts the whole point. Before translating a video, check what on-screen text it contains. If translating is worth doing, it is usually worth rebuilding the few title cards per language — and worth designing future videos with less baked-in text so this stays cheap.
Lip-sync
The top layer: matching lip movement to the new language so the video looks as though it was shot in it. Tools including HeyGen and the avatar platforms offer translation with lip sync, and the result is markedly more natural than dubbing over unchanged footage.
Two things to weigh. It works best on a talking head facing camera at a measured pace, and degrades with fast speech, movement and profile angles. And it produces footage of a real person appearing to speak a language they may not speak — check the publishing platform's rules on synthetic media, label it where required, and think about whether your audience would feel misled on finding out. Being upfront costs nothing next to being discovered.
Quality control
- A native speaker reviews it — for any market that matters. Not the transcript: the finished video, watched through. Tone, register and awkwardness do not show up in text.
- Brief the reviewer properly. "Does this sound like a person from here, is anything unintentionally funny, rude or confusing, and is any term wrong?" A reviewer asked only "is it accurate?" will approve accurate text that reads as stilted.
- Check names and terms against the glossary, including the ones that should not have been translated.
- Cultural context. Jokes, idioms, examples and anything culturally specific need replacing, not converting.
- Watch it end to end before publishing. Every time.
Translation is also a discovery problem
A perfectly localized video that nobody in that market can find has solved half the problem. The half that gets skipped is everything around the video.
- Title and description per language. These are what search and recommendation engines read. An English title on a Spanish-dubbed video tells the platform the video is English.
- Upload the subtitle file rather than only burning captions in. Platforms index subtitle text, so a sidecar file makes the whole video searchable by its content in that language.
- Translate the search terms, do not just translate the words. The phrase people actually type in another market is often not the literal translation of yours — a direct rendering of your title can be a phrase nobody searches.
- Thumbnails with text need their own version per language, for the same reason the in-video graphics do.
Where the platform supports per-language metadata on a single upload, use it — you keep all the engagement on one video instead of splitting it across several. Where it does not, separate uploads per language are usually better than one video with mixed signals, because the recommendation system can then learn who each one is for.
A workflow that does not fall apart at five languages
Doing this once is easy. Doing it every week, across several languages, is where ad-hoc approaches collapse — usually into nobody being sure which version is current.
- Lock the source. Finish the original completely. Every change after this point multiplies by the number of languages.
- Correct the transcript against the glossary, and treat that corrected transcript as the single source everything downstream is generated from.
- Generate per language — subtitles first, then any higher layer you have decided that language deserves. Not every language needs the same layer; your biggest market can have dubbing while the others get subtitles.
- Review, with a named person per language and a deadline. "Someone will look at it" means nobody does.
- Publish with the per-language metadata prepared at the same time, not as an afterthought two days later.
- Record what shipped — which video, which languages, which layer, who reviewed it. When you update the source in six months, this list is what tells you what has to be regenerated.
That last step is the one that separates a pipeline from a pile. Without it, an update to the original quietly leaves four other languages describing the old version, and nobody notices because nobody on the team watches those.
What it costs, and where the money goes
Pricing is usually per minute of audio or video processed, per language, and it climbs steeply with each layer. Subtitles are cheap enough to do routinely; lip-synced dubbing into five languages is a real production budget.
The cost people forget is review. A native-speaker pass is a human hour per language per video, and it is the line item that makes the difference between a translation you can stand behind and machine output you are hoping is fine. Budget it in from the start, or restrict yourself to fewer languages that you can actually check.
The version of this that helps people at home
Everything above is framed around reaching other countries. The same subtitle track does something else that is easy to overlook: it makes your video usable for deaf and hard-of-hearing viewers, and for the very large number of people who watch with the sound off because they are on a train or in an office.
That reframes the cost-benefit of the cheapest layer. Subtitles in your own language are not a translation feature at all — they are an accessibility feature and a reach feature, they cost almost nothing, and they are the first thing to add whether or not you ever translate into another language.
One distinction worth knowing: captions for deaf viewers conventionally include relevant non-speech audio — a phone ringing, music starting, who is speaking when several people are — because those carry meaning that subtitles for a hearing viewer in another language do not need to. If accessibility is the goal, the automatic transcript is a starting point, not the finished article.
When not to bother
- No signal in that market. Translating into a language nobody is arriving from is production spend on a hypothesis.
- Heavily culture-specific content. If the value is local references, examples or regulation, a translation carries the words and loses the point — rewrite for that market instead.
- You cannot get it reviewed. Publishing in a language nobody on your side can read is a public statement you cannot check.
- The original is not working yet. Translating a video nobody watches produces several videos nobody watches.
Common mistakes
- Literal translation. Word-for-word sounds robotic. Localize.
- Skipping review. One embarrassing error travels further than the video did.
- Unreadable subtitles. Too long, too fast, or covering the picture.
- Leaving on-screen text untranslated while the audio is perfect — the giveaway that this is a translation.
- Not fixing the transcript first. Every error you leave at the top is inherited by every language below it.
- Cloning a voice without written consent. Permission for this specific use, in writing, before you generate.
Next step
Understand the voice and avatar technologies behind quality dubbing.