technique
TRACE
techniqueactiveprovisional
trace-98f4854b·1 events·first seen 7d agoAliases: TRACE
Co-occurring entities
More like this (12)
Recent events (1)
TRACE: Tree-structured rollout budget allocation for efficient agentic RL training
TRACE (Tree Rollout Allocation for Contrastive Exploration) is a new framework for improving reinforcement learning with verifiable rewards (RLVR) in multi-turn agentic LLM settings. The method models each ReAct-style thought-action-observation turn as a distinct node, enabling budget allocation across both prompt-level and turn-level prefixes in a tree structure, rather than only at the prompt level. A shared predictor estimates conditional success probability at each anchor to guide allocation, enriching reward contrast within a fixed sampling budget. Empirically, TRACE improves Qwen3-14B multi-hop QA accuracy by 2.8 points over baselines at equal sampling cost.