on-policy-distillation-for-llm-safety-a-routing-approach-to-template-robust-realignment-256a60f5·1 events·first seen Aliases: On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment
A new arXiv preprint introduces Routing-based On-Policy Distillation (ROPD), a safety realignment framework designed to counter malicious fine-tuning attacks that embed harmful behaviors into downstream model corpora. ROPD models the divergence between aligned and compromised output distributions rather than fitting specific prompt templates, addressing three failure modes of existing defenses: catastrophic forgetting, template-mismatch collapse, and re-jailbreaking via system prompt switches. Experiments across four baselines, three datasets, and three base models show ROPD substantially reduces template-mismatch degradation while preserving downstream task performance.