beyond-aggregate-risk-role-stratified-conformal-risk-control-for-llm-tool-calls-7b5925a2·1 events·first seen Aliases: Beyond Aggregate Risk: Role-Stratified Conformal Risk Control for LLM Tool Calls
A new arXiv preprint introduces role-stratified conformal risk control, a calibration method that sets separate statistical risk thresholds for individual semantic argument roles within LLM tool calls (e.g., email body vs. recipient vs. credential). The approach addresses a gap in existing aggregate-level risk certification, where failures in rare high-risk fields can be masked by benign arguments. Evaluated across AgentDojo and InjecAgent benchmarks with six language models, the method achieves consistent per-role budget compliance under model transfer, attack transfer, detector noise, and adaptive attacks. The authors argue that structured tool calls should be certified at the semantic-role level rather than the whole-action level.