HeadlinesBriefing HeadlinesBriefing.com

Targeting Pivotal Decisions for Credit in Agentic RL

Hacker News •
×

Group Relative Policy Optimization (GRPO) has become a promising approach for training large language model agents. However, its uniform assignment of trajectory-level advantages to all policy tokens fails to distinguish consequential decisions from less relevant ones, obscuring which intermediate decisions contributed to success.

We introduce Pro Ver, a framework that targets potentially pivotal decisions for fine-grained credit assignment in agentic reinforcement learning. Given a rollout group, an agentic judge contrasts successful and failed trajectories to propose a segment potentially responsible for their divergent outcomes. Rather than directly trusting the judge's assessment, Pro Ver verifies the proposed segment by estimating its advantage from the difference in terminal success rates between current-policy continuations sampled before and after the segment.

Positive estimates are then incorporated into the GRPO advantages of policy tokens within the proposed segment. By using model judgment only to select where to verify, Pro Ver grounds local credit in observed outcomes without exhaustively evaluating every intermediate state. Across ALFWorld, Web Shop, and Search QA, Pro Ver achieves the strongest average performance at both model scales, with relative improvements over GRPO of 9.91% and 7.12% for Qwen3.5-2B and Qwen3.5-4B, respectively.

Source: Hacker News · Summarized by HeadlinesBriefing