In this guide
Test the section where the wanted sound and the distraction overlap. Leave room around its beginning and end, keep an untouched original, and record the excerpt’s start time before you separate it.
The easiest part of a recording is not always the most useful part to test. An interview introduction may sound clear while traffic covers a sentence later. A quiet verse may separate well while the chorus contains doubled vocals and cymbal wash. Preparing a representative excerpt lets you evaluate the problem you actually need to solve.
This workflow uses Saturalabs’ free audio trimmer and audio file size calculator. Both preparation tools work without an account. AI source separation is a separate service; its results and download credits are handled in the Saturalabs editor.
If your recording has two channels, check whether the wanted microphone or sound is already on one side. The free channel splitter can extract that existing channel without estimating a source. The channels and stems guide explains when this is enough and when a mono mix can lose wanted audio.
1. Choose the part that will decide whether the result is useful
Listen through the original and mark a section that includes your wanted source, the competing sound, and the moment when they overlap. A test containing only a pause can show how quiet the background becomes, but it cannot tell you whether the speaker’s words survive.
| Your task | Include in the excerpt | Listen for after separation |
|---|---|---|
| Clearer interview speech | A full sentence, a pause, and the loudest distraction | Missing consonants, broken phrases, and changes in room tone |
| A karaoke backing track | A chorus with harmonies and a transition into the next section | Remaining syllables, thinned instruments, and uneven transitions |
| An isolated instrument | A note attack, its sustain, and another part playing over it | Lost attacks, interrupted note tails, and bleed from the other part |
| A reusable sound effect | The start of the effect, its full decay, and nearby ambience | An intact initial hit and a natural ending |
For a first listening comparison, a short passage may be enough. Its exact length depends on the sound: a complete spoken thought, musical phrase, or effect matters more than hitting an arbitrary number of seconds. If the recording changes substantially later, make another representative excerpt rather than assuming the first test covers the whole file.
2. Leave context around the cut
Place the start before the first wanted word or note, and the end after the last sound has decayed. Avoid cutting into the attack of a drum hit or the first consonant of speech. Leave enough surrounding ambience to hear whether the separated result starts and ends naturally.
Open the audio trimmer with a mono or stereo MP3 or WAV, set the boundaries in seconds, and prepare a WAV. Its preview plays the same audio that you will download. Listen to the whole excerpt once, then pay attention to the first and last moments. Move the boundaries if the sentence feels chopped off or a reverb tail ends abruptly.
The optional 5 ms fade changes the very edges of the cut. It can soften a cut click; it does not remove clicks elsewhere in the recording. For a sharp sound that begins right at the boundary, move the start earlier or disable the fade and inspect the result. This trimmer keeps one continuous section; it does not join several excerpts.
3. Check file size without unnecessary conversion
Saturalabs’ AI upload accepts MP3, WAV, MP4, and MOV up to 52,428,800 bytes, or 50 MiB. The free trimmer accepts MP3 and WAV recordings up to the same byte size and 5 minutes. For a longer original or a video, create your excerpt in an audio or video editor first.
The trimmer exports 48 kHz, 16-bit PCM WAV. That export can be larger than an MP3 input because it contains uncompressed samples. The editor displays the WAV size and asks for a shorter selection if the export would exceed 50 MiB. It preserves mono or stereo, but does not copy the original file’s tags or artwork.
Use the size calculator to compare duration, channels, sample rate, and bit depth for a planned PCM export. Its fixed-bitrate mode estimates encoded audio payload. Tags, framing, and variable bitrate can make a compressed file differ from that estimate, so inspect the actual exported size before uploading.
For an existing WAV, the free WAV file inspector reads its original sample rate, encoding, bit depth, channels, duration, and actual file size without decoding it. The WAV report guide explains container bits, valid bits, and the limits of header information.
If a clip needs an overall level adjustment for your edit, the free audio normalizer can set a decoded sample peak with one linked gain. It raises or lowers noise along with the wanted sound; it does not make quiet words more even or guarantee better isolation. Read the normalization guide before treating a peak target as a loudness requirement.
When a project or delivery specification requires 44.1 or 48 kHz, the free sample rate converter resamples a supported short mono or stereo MP3/WAV into a 16-bit WAV at that rate. Check the resulting header and timeline alignment. Conversion is for compatibility; it is not required simply to upload a supported recording for isolation, and a higher rate does not restore missing source detail.
Do not raise an export setting to recover detail that is absent from the source. Avoid repeatedly re-encoding a lossy recording just to move it between tools. Keep your best original and make each test excerpt from that same file. If you already have a supported file within the limit, conversion may add work without helping your test.
4. Describe one audible source
After preparing the clip, choose a focused workflow. For dialogue, the voice isolator starts with the speaker you want to keep. For a backing track, the vocal remover starts with the voices you want to separate from the instrumental mix. If you are deciding whether separation is needed, use the audio cleanup method chooser before processing.
Name one audible target, such as “spoken voice,” “vocals,” or “acoustic guitar.” Compare the outputs against the original excerpt. Isolated sound is the requested source; Background audio is what remains. The audio isolation prompt guide has examples for choosing the right output.
Keep your prompt and your selection constant when comparing attempts. Changing both at once makes it difficult to understand what improved the result. Before applying the workflow to more material, inspect the overlap that motivated the test, not only the quieter passages.
5. Keep the excerpt aligned with the original
Write down the excerpt’s original start time. If you cut from 42.50 to 68.00 seconds, the downloaded excerpt begins at its own zero and lasts 25.50 seconds. To place it back in the full edit, start the separated track at 42.50 seconds in the original timeline.
For video, compare lip sync near both ends after importing the WAV. Mute the source audio while auditioning the replacement so the two versions do not play together. Keep a separate copy of the original track for comparison, and review the final video export as well as the audio preview.
If the source clip has been trimmed or moved in your edit, its timeline position may differ from the source time. The video audio replacement guide shows an offset example and how to distinguish a wrong start from drift.
A useful file name records provenance, for example interview-original-start42.50-spoken-voice.wav. Note that the file was estimated from a mixed recording rather than recorded as an isolated source. That gives a future collaborator enough information to trace the edit.
A quick check before processing the rest
- The excerpt includes the wanted sound and the difficult overlap.
- Its beginning and ending preserve complete words, notes, or effect tails.
- The exported file is within the upload limit.
- You have an untouched original, the excerpt’s start time, and the target description.
- The chosen separated output is useful at a comparable listening level, without losing the details your project needs.
If the hardest overlap remains unclear or damaged, try a simpler source description on the original excerpt, or use a different recording if one is available. Trimming makes a test manageable; it does not make missing or distorted audio recoverable.