world-action-drift-attacks-8ca30f41·1 events·first seen Aliases: World-Action Drift Attacks
Researchers introduce BadWAM, a framework for adversarial attacks on World-Action Models (WAMs) — embodied control architectures that jointly predict future world states and actions. The paper demonstrates a new attack class called World-Action Drift Attacks, which use small visual perturbations to break alignment between a model's imagined future and its executed actions, reducing task success from 96.5% to 43.1% in closed-loop evaluation. A stealth variant (imagination-preserving attack) further shows that a WAM can appear to predict a plausible future while executing a desynchronized, harmful action — undermining the safety assumption that world prediction provides a check on action generation.