backdoor-unlearning-generalization-a-path-toward-the-removal-of-unknown-triggers-in-llms-6de7778c·1 events·first seen Aliases: Backdoor Unlearning Generalization: A Path Toward the Removal of Unknown Triggers in LLMs
Researchers demonstrate that training an LLM to unlearn a single backdoor trigger can suppress other backdoors that were never explicitly targeted, a phenomenon they call cross-backdoor transfer. The study spans three model families with backdoors injected via pretraining or continual pretraining, and introduces a new metric called Cross Activation Shift Distance to quantify the relationship between different unlearning interventions. The finding opens a potential defensive strategy where defenders deliberately inject and then remove controlled backdoors to suppress unknown attacker-planted backdoors.