What to fix in your audio before you dub it
Dubbing quality is decided before you upload. Overlapping speakers, music under the voice, reverb and compression degrade it, and none is fixable after.
Artikel ini belum diterjemahkan dan dipaparkan dalam bahasa Inggeris.
A dub can only be as good as the recording it was made from, and most of the quality is decided before you upload anything.
That's not a disclaimer, it's a description of the pipeline. Dubbing has to work out what was said and who said it before it can say the same thing in another language. Anything obscuring that in your source, a second voice, music in the same frequency range, a room with too much reverb, degrades every stage after it.
The part worth knowing: we submit your file as it is. No preprocessing, no cleanup pass, no normalisation. What you upload is what gets dubbed, so anything wrong with it stays wrong.
Why can't it be fixed afterwards?
Because there's no afterwards. Dubbing runs automatically end to end, and nothing is edited by hand once it's done. Our support page says so directly, which means a single bad line can't be re-recorded on request.
That's an unusual thing for a vendor to put in writing, and it's the honest framing for this page. There's no human in the loop who tidies up the output, so the only place you can influence quality is the input.
What actually degrades a dub?
Four things, roughly in order of how much damage they do.
Overlapping speakers. Two people talking at once is the hardest case there is. Where speech overlaps, the transcription underneath has to attribute words to speakers that are physically mixed together, and errors there propagate straight through translation into the output. Our support page's own advice is clean source audio with one speaker at a time, and that's not a formality.
Music sitting in the same frequency range as the voice. Background music isn't automatically a problem. Music occupying the same mid-range band as speech is, because separating them is genuinely ambiguous. A bed that sits low and out of the way survives; a track with prominent vocals or busy mid-range does not.
Heavy reverb. A room with hard surfaces smears each word into the next. Reverb recorded into the file cannot be removed later, and it blurs exactly the boundaries between words that everything downstream needs.
Pre-applied compression and limiting. This one surprises people because it feels like polish. Aggressive compression flattens the dynamic range that distinguishes speech from background, and a heavily limited podcast master is often harder to work with than the raw recording it came from. If you have the pre-master, use that.
What should I do before uploading?
In order of effect:
Record or export the cleanest version you have. If your video's audio was mixed with music and effects, and you still have the voice stem on its own, dub the stem. A separated dialogue track beats a finished mix every time.
Give it one speaker at a time. If you have an interview or a panel, and the speakers are on separate tracks, consider whether the segments can be handled in a way that avoids overlap rather than uploading a mix where two voices collide.
Trim the silence and the non-speech material at the ends. Credits are minutes rounded up with a minimum of one per file, so a trailing 40 seconds of outro music can cost you a whole extra credit for content that carries no speech to dub.
Skip the beautification. Don't add compression, EQ or noise reduction hoping to help. Unless you know the specific problem you're solving, you're more likely to remove information than noise.
Does the file format matter?
Less than the content of the file, but it's worth getting right.
We accept 17 formats: MP3, WAV, AAC, M4A, FLAC, AIFF, OGA, OGG and Opus on the audio side, plus MP4, MOV, AVI, MKV, WebM, M4V, MPEG and MPG for video. The extension is checked before anything else happens.
If you already have a lossless master, upload that rather than an MP3 you exported for distribution. Re-encoding lossy audio to a different lossy format compounds artefacts, and there's no reason to hand the pipeline a worse copy than you have.
If your destination is an audio-only file, such as a YouTube multi-language audio track, upload the audio rather than the video and skip a conversion at the far end.
Limits: 2 GB and 180 minutes per file, both checked before a single credit is touched. Your file's duration is read in memory, straight from the bytes as they arrive, and never written to disk to be measured.
How do I know if it worked?
Listen to the first minute and the last minute, in that order.
The first minute tells you about quality: whether the voice sounds right, whether the translation reads naturally, whether anything is garbled. The last minute tells you about drift, since length differences accumulate across a video and the end is where they show. We've covered why dubbed audio drifts out of sync and what to do about it.
A trial run on one file and one language is the cheapest test available, and if the job fails outright the credits come back automatically, including on the six hour timeout. What isn't recoverable is a job that succeeds but produces a poor dub from poor source audio, because that isn't a failure by any measure the pipeline can apply. It ran, it produced a file, you were charged.
Which is the argument for spending ten minutes on the source before spending credits on the dub.
Where to read next
If you're preparing course content, where clean narration is usually already available and the length difference is the bigger issue, our guide to dubbing e-learning from English to Portuguese covers that pairing, and dubbing SaaS product video from English to Spanish covers the same problem in product demos. The use case index has the rest.
FAQ
Does AI dubbing remove background music?
Don't count on it. Music that sits low and clear of the voice's frequency range generally survives, but music sharing the mid-range with speech makes separation ambiguous and degrades the result. If you have the dialogue stem separate from the mix, dub the stem.
Why does my AI dub sound robotic?
Usually the source rather than the model. Overlapping speakers, heavy reverb and aggressive pre-applied compression all obscure the speech that everything downstream depends on. Flat delivery in the output frequently traces back to a recording that was hard to interpret in the first place.
Should I clean up my audio before uploading?
Remove genuine problems, but don't add processing for its own sake. Use the cleanest version you already have, ideally an unmastered dialogue track, and trim non-speech material from the ends since billing rounds up per file. Adding compression or noise reduction blind tends to remove information rather than noise.
What is the best format to upload?
Whatever is closest to your original. We accept 17 formats and submit the file as it is with no preprocessing, so a lossless master is better than a distribution MP3. Re-encoding one lossy format to another only compounds the artefacts.