self-state-attacks-on-self-hosted-ai-agents-how-far-can-os-defenses-go--56ddf9fa·1 events·first seen Aliases: Self-State Attacks on Self-Hosted AI Agents: How Far Can OS Defenses Go?
A new arXiv preprint introduces 'self-state attacks,' a threat class where AI agents are compromised through corruption of their own memory and configuration files via legitimate OS system calls. The authors formalize a four-axis attack space, instantiate it as a 23-cell matrix with 43 concrete operations on real agent state files, and evaluate layered OS defense strategies including access control, workload-conditioned detection, and periodic backup. Results show a layered defense stack handles most attack cells but a residual attack surface remains structurally indistinguishable at the OS level, suggesting OS-level defenses need rethinking for agentic deployments.