You could identify the song and then subtract it. If you add which noise it raises the noise floor, but won't effect overall decoding that much, because there is such deep knowledge of which phonemes are likely to follow in a given sequence. You have to assume there is a model trained for each participant, e.g. using telephone intercepts, other listening devices.
Maybe if you simultaneously played back segments of dozens of conversations of the participants talking. That would certainly be confusing for the participants.
Maybe if you simultaneously played back segments of dozens of conversations of the participants talking. That would certainly be confusing for the participants.