sovereignpa-bench-5287b89a·1 events·first seen Aliases: SovereignPA-Bench
Researchers introduce SovereignPA-Bench, an executable benchmark designed to evaluate personal AI agents on dimensions beyond task completion, including privacy preservation, consent compliance, evidence grounding, manipulation resistance, and user burden. The benchmark tests 120 sovereignty stress scenarios across 4 model families and 8 policy baselines, generating 3,840 frozen-prompt trajectories with a blinded human audit component. Results show that full-sovereign scaffolding outperforms all baselines on sovereignty metrics while reducing privacy leakage and consent violations. The work argues that personal-agent evaluation must shift from task success toward consent-aware, evidence-grounded action.