Localize video instead of translating it
Most video localization advice is really about translation: take the finished file, replace the words, add subtitles, line the new voice up with the old mouth. That produces a video that is in the right language and still looks dubbed. BrandU keeps what made the video work and rebuilds it in the target language, so the picture and the language are made together rather than stitched together afterwards.
By Elias Sun, Founder of BrandU
Why dubbed video looks dubbed
There is a reason the phrase “it looks dubbed” exists, and it is not the quality of the voice recording. It is that the video was made for one language and then had another one put on top.
The mouth was shaped for different sounds. The sentence lengths were set by different grammar. The on-screen text was written for a different reader. A translation workflow can align all of that — that is what lip sync, timing shifts and subtitle timing are for — but alignment is a repair, and repairs show.
None of that is a failure of the translation. It is what happens when the language changes and the video does not. It is also why dubbing services and voice over translation put so much of their effort into alignment rather than into performance: the performance was already fixed by the original shoot, so the only lever left is how well the new words can be bent to fit it.
The difficulty there is inherited, not chosen. It exists because the video and the language were produced separately — so the only lever left is alignment. Produce them together and there is nothing left to align. That is the whole difference, and it is what makes a localized video look made for its market rather than adapted to it.
So the answer is not a better translation. It is a video that was built in the target language in the first place — same structure, same pacing logic, same argument, new language, and a picture that was made for those words.
| What the original fixed | What translating has to work around |
|---|---|
| Mouth shapes for its own sounds | New sounds that do not match those shapes |
| Segment lengths set by its grammar | Translations that run longer or shorter |
| On-screen text written for its reader | Text that now has to be replaced |
| A cultural reference the audience knows | A reference that now needs explaining or swapping |
What actually changes per market
Localize is often used as a synonym for translate, and the two are not the same thing. The difference shows up the moment you list what actually has to change.
Two rows in that list are the reason a translated video can be perfectly accurate and still land wrong: the language was localizable, the video was not.
This is what makes multilingual marketing harder than it sounds. You are not shipping one asset in five languages. You are shipping five versions of an idea, and the idea has to be re-expressed each time — which is a production question, not a translation question.
Ad localization is where this bites first, because ads carry more than language. An offer, a currency, a season and a claim all have to be right for the market, and all four of them can be wrong while the translation is perfect.
| Layer | What it is | Does translating it work? |
|---|---|---|
| Language | The script and the voice | Yes, that is what translation does |
| On-screen text | Captions, labels, price and offer | Only if it is replaced, not covered |
| Reference points | A holiday, a norm, a unit, a currency | Rarely — it usually needs rewriting |
| Casting and context | Who is on screen, where, wearing what | No — that is the picture |
| Placement and ratio | Where it runs, in what shape | No — that is the format |
One master, every market
Video localization services usually split into two offers: a translation pipeline, or a reshoot in each market. The first is cheap and shows its seams. The second is clean and costs a production per language.
There is a third position, and it is the one this tool takes: the structure is the master, and the language is one of its variables.
One master. Every market is an instance of it. Which is what changes the economics: the second market is not a second production, and the fifth is not a fifth.
There is a second effect that shows up later and matters more. When a market is a production, teams stop adding markets — not because a new market would not pay off, but because each one adds a permanent production commitment. When a market is an instance, the calculation flips: the marginal market is cheap enough that the question becomes whether the market is worth having, which is the question you actually want to be answering.
Which markets first: the usual advice is to start with the largest market for your category. That is a reasonable default and a poor one to follow blindly, because the cost of a market is not the same for all of them. Two things make a market cheap to add — how much of the video survives, and how stable the market's commercial details are. A market whose offer changes every month is a market you will re-edit every month, which is fine when that is a segment change and expensive when it is a new localization project.
| What you keep | What you change per market |
|---|---|
| The order of the argument | The script language |
| How shots connect and where they cut | The voice and read |
| Caption placement and pacing | The captions themselves |
| The overall format and feel | Ratio, where the placement needs it |
Per market, not per language
Two markets can share a language and still need different videos. The UK and the US both read English; the offer, the currency, the spelling and the reference points do not match. Treating localization as a language list misses that.
That matters most where the volume is: ecommerce localization is not five translations of one product video. It is the same catalogue, adapted to each market's norms, currency, season and shipping promise — and then kept current when any of those move.
The useful distinction is the last row of that comparison. A paragraph can be translated and dropped in. A person speaking on camera cannot — unless the video was built so that the speaking part is a variable.
That is worth saying plainly, because content localization has a reputation for being a spreadsheet exercise: a list of strings, a list of markets, a column you fill in. It is that, for text. Video has no column you fill in — the equivalent of that column is a new take, and whether you can afford one is the whole question.
| The assumption | What people usually do about it | What actually happens |
|---|---|---|
| One language = one version | Translate once, publish everywhere | UK and US English still diverge on offer, currency and reference |
| Localization is a one-time project | Run it as a project with a delivery date | Prices, seasons and promotions keep moving, so it is never finished |
| Content localization is the same as video localization | Keep a spreadsheet of strings per market | A string can be swapped in place; a performance cannot |
What stays, what changes
This is what rebuilding means here, component by component.
The left column is the answer to what gets kept: structure, not pixels. The right column is the answer to what gets rebuilt: everything a viewer in that market would notice. Because the speaking parts are variables, the mouth is not a problem to solve — it is generated for the words it is saying.
There is a mechanical reason this works, and it is worth one sentence: because the video is built as independent segments rather than one flat file, changing the language re-times the rest. The timeline and the related elements follow on their own, so you never adjust a timeline by hand.
In practice that turns localization from a project into a setting. A project has a kickoff, a scope and a delivery date. A setting has a value you change when the market changes. The reason teams treat localization as a project is not that it should be one — it is that in a translation pipeline, the only way to change the language is to start another project.
| Component | Kept — the part that sells | Replaced — for this market |
|---|---|---|
| Script structure | The order of the argument | The language it is argued in |
| Shots and scenes | How shots connect | Your footage and product |
| Captions | When and where they appear | The translated copy |
| Voiceover | Tone and rhythm | The voice, in the new language |
| B-roll | Where it cuts in | Your material |
| Effects | How they appear | Your brand feel |
| Pacing | Information density | What you emphasise |
| Format | Vertical ratio | Your platform |
What a change costs
Five common edits on one 30-second video, and what each costs.
| What you do | Without reuse | With BrandU |
|---|---|---|
| Create the first video | — | 1,030 |
| Change the order of segments | 1,030 | 350 |
| Change an element's styling | 1,030 | 690 |
| Swap in different assets | 1,030 | 180 |
| Rewrite one segment's copy | 1,030 | 180 |
Costs are shown in credits. Values derive from the same pricing contract that powers the pricing page.
What this looks like in practice
One master, five markets, no reshoot.
A DTC brand selling into five markets
- Channel
- Paid social, one placement per market
- Format
- One proven format, rebuilt per market
- Workflow
- Keep the master; each market is a language and an offer, not a new shoot
- Outcome
- Five market versions live, and a price change is a segment edit
What BrandU does
Four things, in the order they matter when you are adding a market:
- It rebuilds the video in the target language rather than translating it. The script is written in the new language, so the voice-over is native rather than aligned, and the picture is generated for those words — which is the part an AI video translation workflow has to work around, because translation starts from a finished file.
- It keeps one structure as the master. Every market you add is an instance of it: the argument, the shot logic, the caption placement and the pacing carry over, and the language is the variable.
- It makes a new market an edit, not a production. Changing the language re-runs the affected segments rather than rebuilding the video, which is why the fifth market is not the fifth shoot.
- It keeps markets current. Prices, offers and seasons move; a change to one of them is a segment change, not a new localization project.
- It puts the effort into the message rather than the mechanics. Rewriting a line, changing an offer or switching a market is an edit in the script, so the hours go into what the video argues in that market — not into re-cutting, re-timing and re-exporting it.
How to run this
You do not have to start from a reference video. If you have one, that is the strongest path, because the structure is already proven. If you do not, the create screen also takes a one-sentence description or a template — three ways in, same editor afterwards.
- Find a video that is already working in the category. Paste its page link — no download, no direct file URL needed.
- Bring your brand and product information in once. It is reused across every market.
- Pick the markets. A market is not the same thing as a language — set the offer, currency and reference points for each.
- Generate each version. The structure stays; the language and the speaking parts are new.
- When a price or an offer moves, change that segment. Do not redo the localization.
FAQ
- Video localization is adapting a video so it works in a specific market — not just in that market's language, but for its norms, currency, offer and reference points. Translation is one part of it. The part people underestimate is that the speaking parts and the on-screen text have to be remade, not overlaid, or the result reads as dubbed no matter how good the translation is.
- No, and the difference is the whole point. Translation changes the words. Localization changes the video. A translated video keeps the original picture and puts new language on top, which is why it needs lip sync, timing shifts and subtitle work to look acceptable. A localized video is built in the target language, so the picture and the words were made for each other.
- Every language the video needs to be made in. The script is written in the target language rather than translated into it, so the language you publish in is the language you write in.
- No — and that is the difference from the usual two options. A translation pipeline avoids the shoot and shows its seams; a reshoot per market is clean and costs a production each time. Here the structure is the master and the language is a variable, so a new market is an edit rather than a production.
- They solve the language, and they solve it well. What they cannot change is that the picture was shot for a different set of sounds and sentence lengths — so the work goes into alignment. Building the version in the target language removes the alignment problem instead of solving it.
- It produces short-form video, which is where the volume and the testing pressure sit. What carries across is the shape of the deliverable: one structure kept as the master, with language and speaking parts as variables. That is the same shape that makes drama dubbing cheaper to do properly — a serialised format is exactly where adapting market by market gets expensive, because there is so much of it and the language is the whole product.
What is video localization?
Is localization the same as translation?
How many languages does this support?
Do I need to reshoot for each market?
What about dubbing services and voice over translation?
Does this cover drama dubbing and longer content?
See related pages: Video ad use cases · How cloning a video works · Ecommerce video ads · Creative production · Content ideas for video ads
See BrandU in your own video
Clone a video that already sells, rewrite the script with your product, render in minutes. You only pay for what changes.