AI video translator

Translate a video into another language, dubbed and captioned

Drop in a video and get it back in another language: transcribed, translated as natural speech, read by a new voice over your picture, with captions in the new language burned in. Dub voices cover 14 languages.

Everything on this page was made on Treza.

Whisper reads the original

The source is transcribed with Whisper, which is multilingual and detects the spoken language on its own. That transcript is the script the rest of the dub works from, so nothing has to be typed up by hand.

Translated to be spoken, not read

A language model rewrites the transcript as narration a native speaker would actually say, never word for word. It keeps names, brands, and numbers accurate and keeps the length close to the original, so the dub fits the video.

14 dub languages

Dub voices speak English, Spanish, French, German, Italian, Portuguese, Dutch, Polish, Russian, Chinese, Japanese, Korean, Hindi, and Arabic. Each dub voice is labelled with its language in the picker, so matching the voice to the target is quick.

Mixed over the original, broadcast style

The new voice plays over your picture while the original soundtrack ducks far underneath it, then comes back up between lines, so the music and room tone of the original survive. The frame size and frame rate match the source.

Captions in the new language

The finished dub is transcribed again on its own timeline, with the written translation as a spelling guide. The burned-in captions are in the target language, spelled the way the translation wrote them, and land when the new voice says them.

A check before anything is burned in

If less than half of the spoken dub matches the written translation, the caption step stops the run rather than burning in the wrong words. The captions you publish are the words the voice said.

Every stage is a node you can change

Tell the translator to keep a formal register or leave product names in English, swap the voice, or restyle the captions. Lock the output of a step you are happy with and the next run redoes only the rest.

What a dub is made of

A transcript, a new voice mixed over the picture, and captions timed to that voice. Each tile is a real run that took one of those steps, credited to the model or template that made it.

A mission webcast clipped to vertical from its own transcript

It starts with a transcript

Whisper reads the source, the same model the dub uses.

Whisper Large v3

A fact of the day short, narrated over four generated scenes

A new voice over the picture

Narration mixed over the shots, the rest ducked under it.

Fact of the Day

A captioned short with karaoke-style word highlighting

Captions timed to that voice

Word timings from the finished audio drive the highlight.

Captioned AI Short

From idea to finished video

  1. Step 01

    Drop in the video

    Upload a file to the Source Video node or paste a direct MP4 link. Any video with speech works: a tutorial, a product demo, an interview, a clip for the feed.

  2. Step 02

    Name the language and pick a voice

    Type the target language into the Target Language node and choose a Dub Voice that speaks it. The template starts on Spanish with a Spanish voice, so a first run works as it is.

  3. Step 03

    Run it

    Whisper transcribes the original, the translator writes the spoken version, the voice reads it, and a Sequence node mixes it over your picture with the original ducked underneath.

  4. Step 04

    Check the dub and ship it

    Preview the dubbed video with its translated captions on the canvas, then download the MP4 or wire it into the YouTube or TikTok upload node. For the next market, change the language and the voice and run it again.

The Treza node editor showing a scheduled daily short pipeline running left to right: a brief and a schedule into a writer model, then a Seedance 2.5 video node, transcription, burned-in captions, and a YouTube upload node

Ask for the pipeline in the composer, or send it from Claude over MCP.

What people build with it

Courses and tutorials for a new market

A lesson recorded once, in your language, becomes a Spanish, Japanese, or German version with captions, without recording anything again.

Product demos in the buyer's language

Walkthroughs and explainers read in the language of the market you sell to, with names and numbers kept exactly as the original said them.

A second-language YouTube channel

Dub each upload for a Spanish or Portuguese audience and publish it as its own video, from the same source file you already have.

Shorts and social clips in several languages

Short clips are the sweet spot: one continuous read over the whole clip, captioned in the new language for muted autoplay.

Interviews and talks, voice-over style

The way broadcasters translate a guest: the translated voice carries the meaning while the original ducks underneath and returns between lines.

Training and company updates

Record one update and send it to teams in several countries, each hearing and reading it in their own language.

One video. A version for every audience.

Change the language, pick a voice, and run it again.

AI video translator, answered

How does the AI video translator work?

It is the Video Dubbing template, one pipeline on one canvas. Whisper transcribes your video, a language model translates the transcript into natural spoken narration, a text to speech voice reads it, and a Sequence node mixes that voice over your original picture with the original soundtrack ducked underneath. The dub is then transcribed again so captions in the new language are burned in, timed to the new voice.

Which languages can I translate a video into?

The translation step takes any language you type. For the spoken dub you need a voice that speaks it, and the multilingual voice model covers 14 languages: English, Spanish, French, German, Italian, Portuguese, Dutch, Polish, Russian, Chinese, Japanese, Korean, Hindi, and Arabic. On the source side, Whisper is multilingual and handles major languages well, with English strongest.

Does it sync the speaker's lips to the new language?

The template dubs voice-over style, the way broadcast translation works: the new voice plays over the picture and the mouth stays as filmed. When the speaker is on camera and you want the mouth to follow the new words, add a Lipsync node, which reshapes a speaker's mouth to match whatever audio track you wire into it.

Whose voice reads the dub?

A voice you choose from the text to speech catalog, matched to the target language, and you can preview each one on the canvas before a run. The translation keeps the meaning, tone, and energy of the original. Pick the one closest in character to your speaker.

How long can the video be?

It does its best work on videos up to a few minutes long. The dub is one continuous read rather than a line-by-line match, so across a long video it drifts from the original speech timing. For longer material, cut it into sections with an Extract Clip node and dub each one.

Can I control how it translates?

Yes. The translator is a language model node with its instructions in plain view, running on Claude Sonnet 5 by default. Ask for a formal or casual register, give it a list of product terms to leave untranslated, or switch it to another model. Out of the box it keeps names, brands, and numbers exactly as the original said them.

What do I get back?

An MP4 of your original picture with the dubbed voice mixed in and translated captions burned in, at the source's frame size and frame rate. The transcript, the translation, and the voice track are node outputs too, so you can take any one of them on its own.

How much does it cost to make ai video translator with Treza?

Treza runs on prepaid credits with no subscription. Each generation is charged at the model's own rate from your balance, and only successful runs are charged, so a failed generation costs nothing. A typical video generation settles around $1.06, and credit packs start at $5 and never expire.

Can I call this as an API instead of using the canvas?

Yes. Every pipeline can be published as a versioned HTTP endpoint. Call the typed /invoke endpoint with JSON in and JSON out, or point any OpenAI SDK at the OpenAI-compatible endpoint. Swap a model on a node later and the API your product calls does not change.

Do I need to know how to edit video?

No. Start from a template, change the topic, and run it. If you do want frame-level control, the timeline editor is there with multi-track video and audio, per-clip trims, fades, and volume, but nothing about the automated path requires opening it.