In this guide
Name one audible source: “spoken voice,” “acoustic guitar,” or “dog barking.” Then choose Isolated sound to keep it or Background audio to hear the rest of the recording.
A text-guided audio separator needs a target. “Make this sound better” describes your goal, but it does not identify a sound in the recording. A more useful starting point is a short description of something you can hear.
The descriptions below are examples to test, not a list of commands with guaranteed outcomes. Separation quality depends on the recording and how clearly the sources can be distinguished.
Start with a source, then add a useful detail
Use a noun or short phrase first: “vocals,” “drums,” “spoken voice,” or “door closing.” If that target is too broad, add one characteristic that distinguishes the sound you mean, such as “acoustic guitar” or “lead singing voice.”
Avoid piling on quality adjectives. “Perfect, professional, crystal-clear audio” does not tell the separator whether to extract dialogue, music, or an instrument. Likewise, mentioning several unwanted sounds can make your intention less clear than naming one wanted source.
Choose which output you actually need
Saturalabs separates a described target from the rest of a recording. This creates two useful ways to approach the same edit.
| Your goal | Example target | Output to inspect |
|---|---|---|
| Keep an interview voice | Spoken voice | Isolated sound |
| Make a karaoke backing track | Vocals | Background audio |
| Extract a guitar part | Acoustic guitar | Isolated sound |
| Reduce barking in a clip | Dog barking | Background audio |
| Collect a sound effect | Door closing | Isolated sound |
Both strategies need listening checks. Isolating a distracting sound and keeping the remainder may remove some wanted sound too. Isolating a wanted voice may leave some background attached to it. The better result is the one that works in your finished project.
Run a small, repeatable comparison
- Pick a representative section. Include the target and the sounds that overlap it. A quiet introduction may not reveal problems in the busiest part.
- Write down the first target. Start with a simple phrase and keep the original file unchanged.
- Listen to both outputs. Check whether the wanted source appears in the isolated track and whether recognizable pieces are left in the background.
- Change one detail. Compare “guitar” with “acoustic guitar,” for example, rather than rewriting the entire prompt.
- Decide using the intended use. A practice reference, a dialogue edit, and an exposed sample have different tolerance for artifacts.
Record a short note for each attempt: prompt, section tested, what improved, and what was damaged. This makes it easier to choose a result after several attempts instead of relying on memory.
Diagnose an unhelpful result
The wrong sound was extracted
Check the description and the selected output first. If you asked for “vocals” and want an instrumental, the background track is the relevant one. If you wanted speech but also extracted singing, try “spoken voice” and compare the same section again.
Only part of the target is present
Listen where the target changes character or another source becomes louder. Try a broader description if your first one was overly specific. If the missing part is already hard to hear in the original, a different prompt may not recover it reliably.
Two similar sources remain together
Two singers or two guitars can overlap in ways that are difficult to separate from a finished mix. Do not assume that adding more words will resolve this. If you can obtain the original tracks, those provide a better starting point for editing individual performers.
Finish with a listening checklist
Listen at a consistent volume. Check a quiet section, a busy section, and the transition between them. Compare the processed sound by itself and in the project where you intend to use it.
- Is the target recognizable throughout the section?
- Are any words, notes, or transient details missing?
- Is background bleed acceptable in context?
- Does the result introduce a distracting metallic or watery texture?
- Would a simpler edit solve the problem with less damage?
Choose a tool that matches the source you can hear: the voice isolator starts with a speaking voice, while the sound effect extractor gives examples for distinct events such as bells and impacts. Both use the same target-versus-remainder output choice.
For an applied example, follow the dialogue cleanup guide. For musical material, the karaoke guide shows which output to choose and how to check a backing track.