In this guide
Keep the picture in place. Add the replacement audio at the matching source position, mute the old sound, and check sync near both ends. A cleaned excerpt belongs at its original offset, not automatically at the start of the video.
To replace audio in a video, add the new recording to an audio track, align it with the matching picture, and mute the sound it replaces. Export the finished timeline with audio enabled. The important part is keeping the relationship between the original clip and its replacement intact.
This guide covers a full-length WAV and a shorter excerpt returned from an audio cleanup tool. Saturalabs accepts supported MP4 or MOV uploads, but its separated downloads are audio files. Import your chosen WAV into an editor to create the finished video.
The short workflow
- Keep the original. Duplicate the project or timeline before replacing any sound.
- Choose the right audio. Audition the WAV and confirm it contains the voice, music, or remaining mix you intend to use.
- Match the position. Put the new audio at the same source moment as the picture. Retain any excerpt offset.
- Mute the replaced sound. Listen to the replacement by itself, then with the other sounds you deliberately want in the edit.
- Check the export. Review the beginning, middle, and end of the rendered video, including transitions and the last spoken word.
Choose the WAV you actually want
If you used the video dialogue cleanup workflow with a prompt such as “spoken voice,” inspect Isolated sound. That is the estimated voice. If you targeted an unwanted sound, inspect Background audio: it contains what remains after the requested target is separated.
For example, when you target music in a video soundtrack, Background audio is the candidate for a version with that music reduced. A prompt naming the speaker instead produces a different candidate: the isolated voice may also omit ambience or other sounds you wanted. Neither output is a guarantee that every wanted detail survived.
Compare a complete sentence and the hardest overlap with the original before downloading. Review the account and credit details shown for the selected download. Save the file with a useful name that records the target and excerpt position; “interview-start42.50-voice.wav” tells you more than “final.wav.”
Replace audio in a video editor
Use the project where the video is already edited. Import the WAV and put it on its own audio track beneath the matching clip. Keep the original audio available on a muted track so you can compare and undo the replacement.
If selecting audio also selects its picture, use your editor’s unlink, detach, or audio-only selection function before changing the sound. Check the selected items before deleting anything. In Premiere, Adobe’s linking documentation explains Clip > Unlink for independent editing and Clip > Link to move matched audio and video together.
- Place the WAV at the matching start position, then zoom in to inspect a recognizable word, clap, or other transient.
- Mute the original clip audio. If its track contains other clips that should still play, mute only the replaced section rather than the whole track.
- Listen across the replacement boundaries. Preserve complete words and add restrained fades or suitable room tone where the edit needs continuity.
- After confirming alignment, link or group the replacement with its picture if your editor supports that. Keep track of the reference audio separately.
- Recheck sync after moving clips, changing speed, or making ripple edits. A link does not determine the correct initial position for you.
Do not delete the replacement’s quiet beginning simply because its waveform looks empty. That interval may be the timing handle that keeps later speech aligned. Trim it only as part of a matching timeline edit.
Place a trimmed WAV at the right time
A WAV made from an excerpt starts at its own zero. It does not automatically remember where that excerpt belonged in the source video. Record the source start and end when preparing the clip; the audio preparation guide explains that handoff.
| What you processed | Where the picture sits | Where the WAV starts |
|---|---|---|
| The full source, beginning at 0 seconds | Source begins at timeline 0 | Timeline 0, followed by a sync check |
| Source seconds 42.50–68.00 | Full source begins at timeline 0 | Timeline 42.50; the excerpt lasts 25.50 seconds |
| Source seconds 42.50–68.00 | Source second 20 begins at timeline 100 | Timeline 122.50: 100 + (42.50 − 20) |
For one continuous clip at normal speed, the placement is: clip timeline start + (excerpt source start − clip source in). Use seconds for all three values; a frame-based timecode is a different representation. The excerpt must fall within the source range that the clip actually uses.
This arithmetic does not cover a montage, speed change, freeze frame, or time remapping. If the picture uses several source ranges, match the audio to those ranges separately. Often the simplest handoff is to export the exact edited section with its current timing, process that section, and return the result to the same timeline range.
Find doubled audio, constant offset, and drift
Check a clear visual sound cue near the beginning and another near the end. A visible clap is useful; for dialogue, inspect a sharp consonant and watch the mouth movement while listening. A changed waveform shape after separation makes a single visual peak less dependable than several cues together.
| What you notice | What to inspect first | A useful next check |
|---|---|---|
| Echo, hollow tone, or unexpectedly loud speech | Original and replacement may both be playing | Solo the replacement and inspect clip and track mute states |
| The same timing error at both ends | Starting position or a trimmed leading interval | Move the audio by the measured offset, then verify both cues again |
| Beginning matches, ending does not | Duration, speed changes, or how the source was interpreted | Compare the actual source ranges and durations before sliding the start |
| Sync jumps after a cut | A source discontinuity or timeline edit | Match each edited range; one offset cannot correct several different cuts |
| Words vanish even when timing is right | The selected separation result | Return to the original and compare Isolated sound and Background audio |
A growing error needs more investigation than moving the entire WAV. A wrong start and a wrong duration are different problems. Avoid stretching speech just to make the last word fit before you have checked whether you exported the right section.
Optional: replace full-length audio with FFmpeg
For an already aligned, same-duration WAV and a compatible MP4, this command creates a new file. Run it in the folder containing the two inputs, after installing FFmpeg. Replace the example file names with yours.
ffmpeg -n -i "original.mp4" -i "cleaned.wav" -map 0:v:0 -map 1:a:0 -c:v copy -c:a aac -b:a 192k "replaced-audio.mp4"
The maps select the first video stream from the original and the first audio stream from the WAV. Video is copied; audio is encoded as AAC at a requested 192 kbit/s. The original soundtrack and subtitle streams are not selected. -n refuses to overwrite an existing output. See FFmpeg’s stream copying and mapping documentation.
This example does not align an excerpt, mix tracks, repair drift, or match unequal durations. It deliberately omits -shortest, which ends output at the shorter stream. Use a timeline editor for mismatched lengths or offsets. Stream copying also requires video compatible with the MP4 container.
We verified the published command with generated four-second H.264 clips and distinct test tones using FFmpeg 8.0.1. The output kept the original video packet hashes and timestamps, contained the replacement tone rather than the original tone, and refused to overwrite an existing file. A two-second WAV left the four-second video intact but did not supply sound for the missing interval. These fixture checks verify file assembly, not cleanup quality or every input codec.
Check the finished video before sharing it
Export a short representative section first. Keep the intended picture framing and project timing, enable the desired audio output, and check that the replacement track reaches the final mix. A WAV preview cannot confirm what your video export includes.
- Listen without the original track competing with the replacement.
- Check sync near both ends and on either side of edits.
- Confirm that soft words, breaths, effect tails, and the last sentence remain audible.
- Check music and ambience deliberately: voice isolation may remove them along with the distraction.
- Open the rendered video in a player and review its duration and sound before replacing or sharing any delivery file.
If the replacement is less intelligible than the original, return to the source rather than hiding the damage with louder playback. Use the cleanup method comparison to consider a different approach, and keep the original project available.
Common questions
Does Saturalabs download a video with its audio already replaced?
No. Separated downloads are WAV audio. Add the selected WAV to your video editor, align it with the picture, and export the finished video there.
Why does the cleaned voice sound doubled?
Check whether the original soundtrack and replacement are both audible. Solo the new track, then mute only the original audio it replaces. Keep unrelated music or effects if the edit needs them.
Where do I place audio made from a short excerpt?
Place it at the corresponding source position in your timeline. If source seconds 42.50–68.00 were processed and the full source starts at timeline zero, the WAV starts at timeline 42.50, not zero.
Will changing the sample rate automatically fix sync?
No. A correctly converted file should retain its duration. Check the source range, starting offset, and clip speed first. Changing a rate without understanding how the file is interpreted can introduce another timing error.
Can I replace the soundtrack without re-encoding the video?
Yes, when the original video stream fits the output container and your operation does not need picture changes. The FFmpeg example copies video and encodes the replacement audio; it is intended for audio already aligned to the full clip.
