Your video, speaking a language you don’t

This AI video translator transcribes the speech in your video, translates it, generates a voiceover in the target language, and lip-syncs it back onto the footage. Eleven languages are supported end to end, and the source language is detected automatically unless you name it.

How it works

  1. 1

    Point it at a video

    A fetchable video URL — the lip-sync step reads it directly.

  2. 2

    Choose the target language

    One of eleven. Optionally name the source language instead of auto-detecting.

  3. 3

    Dub and sync

    Transcribe, translate, voice, then lip-sync back onto the original footage.

What you can make

Indonesian creators going global

Publish the same video in English without re-recording, and keep your own face and delivery.

Global brands localising

One campaign video becomes several markets without booking voice talent per language.

Course and training teams

Translate a module once and keep the instructor on screen, which matters more for retention than subtitles.

Anyone with a back catalogue

Old videos that performed in one language are the cheapest content you own to test in another.

This AI video translator transcribes the speech in your video, translates it, generates a voiceover in the target language, and lip-syncs it back onto the footage.

Dub your first video free

Specs & limits

InputA fetchable video URL (or a video from your library)
Target languagesEnglish, Bahasa Indonesia, Spanish, French, German, Portuguese, Italian, Japanese, Korean, Chinese (Mandarin), Hindi
Source languageAuto-detected, or name it yourself
What happensTranscribe → translate → voice → lip-sync back onto the video
BillingPer second of video — the lip-sync step dominates the cost
Commercial useYours to use, including client work

Tips for better results

  • Cost scales with video length, not word count, because lip-sync is billed per second. Trim to the section you actually need translated before running it.
  • Use clean source audio. Everything downstream depends on the transcript, so background music under speech costs you accuracy in the translation as well as the voiceover.
  • Name the source language when the speaker code-switches. Auto-detect settles on one language, and a video mixing Indonesian and English can be read as the wrong one.
  • Expect the translated audio to run a different length. Some languages need noticeably more syllables for the same meaning, and pacing shifts even though the lip-sync holds.
  • Check the first sentence before committing a long video. If a term of art came through wrong there, it will be wrong throughout.

Frequently asked questions

Which languages are supported?

Eleven, end to end: English, Bahasa Indonesia, Spanish, French, German, Portuguese, Italian, Japanese, Korean, Chinese (Mandarin) and Hindi. The list is deliberately curated — these are the ones verified through the whole transcribe-translate-voice-lipsync chain, rather than a longer list that only half works.

Does it move the lips?

Yes — the translated voiceover is lip-synced back onto the original footage, so the speaker appears to be saying the new words rather than having audio laid over them.

Do I need to say what language the video is in?

No, it is detected automatically. Naming it is worth doing when the speaker mixes languages, since auto-detect has to settle on one.

How is it charged?

Per second of video. Transcription and the voiceover add on top and vary with how much speech there is, but the per-second lip-sync rate is the dominant cost — so a shorter clip is dramatically cheaper.

Is it free?

Generating uses credits from your Publiz plan. This is one of the more expensive tools because of the per-second lip-sync; new accounts start with a free credit allowance you can test a short clip on.

Related tools

Browse all AI tools