Run two songs through the same vocal removal and one comes out spotless while the other keeps a faint version of the original singer hanging behind the band. The difference is not the software having a bad day. It is the mix, and once you know which features make a song hard you can usually predict the result before you press anything.
The old approach to this problem was arithmetic. Lead vocals usually sit in the centre of a stereo mix, so subtracting one channel from the other cancels anything in the middle. It does remove the voice. It also removes the bass, the kick drum and the snare, because those are in the middle too, and it collapses the song to mono on the way past. More on why that trick disappoints.
What modern separation does is different in kind. A model has listened to an enormous number of songs alongside their separated parts, and has learned what a human voice looks like as sound: how its harmonics stack, how it slides between notes, how it breathes, how it differs from a saxophone playing the same line. Given a new mix it estimates which part of what it is hearing is voice, and pulls that out.
So it is not cutting a frequency range, and it is not cancelling a position in the stereo field. It is making a judgement, thousands of times a second. And like any judgement, it is confident in obvious cases and hesitant in ambiguous ones. Every hard case below is a case where the model has genuine reason to be unsure.
A dry voice is contained. It starts, it stops, and while it is happening it is clearly in one place. Reverb takes that voice and smears a decaying copy of it across the next second or two of the song, mixed in with everything else.
That tail is still voice, so the model would like to remove it. But by the time it is a tail it is quiet, spread across the whole spectrum, and tangled up with the drums and guitars that happened after the note ended. Remove it too eagerly and you take the room off the whole record and the band sounds like it is playing in a cupboard. Leave it and you get the ghost: no words you can make out, but a wash that is unmistakably the shape of a voice that was there.
This is why big ballads and a lot of older recordings, where the reverb is part of the sound rather than an effect on top of it, are consistently the hardest. It also explains why the residue you hear is usually vague rather than intelligible. What survives is the tail, not the word.
Useful to know before you decide a song has failed:
Judge it with the song playing at the volume you will actually sing at, not on headphones at low volume hunting for flaws. A residue you have to strain to notice is not going to be there when you are singing.
And if a song is important enough to you that none of this is good enough, it is worth remembering that the hard cases are hard for everyone. There is no setting anywhere that separates a heavily reverbed ballad perfectly, because the information needed to do it cleanly is not in the recording.
Free public beta, on any iPhone running iOS 17 or later. No account.
Join the free beta