Agentic Rubrics as Contextual Verifiers for SWE Agents
Mohit Raghavendra, Anisha Gunjal, Bing Liu, Yunzhong He
摘要
Verification is critical for improving agents: it provides the reward signal for Reinforcement Learning and enables inference-time gains through Test-Time Scaling (TTS). Despite its importance, verification in software engineering (SWE) agent settings often relies on code execution, which can be difficult to scale due to environment setup overhead. Scalable alternatives such as patch classifiers and heuristic methods exist, but they are less grounded in codebase context and harder to interpret. To this end, we explore Agentic Rubrics: an expert agent interacts with the repository to create a context-grounded rubric checklist, and candidate patches are then scored against it without requiring test execution. On SWE-Bench Verified under parallel TTS evaluation, Agentic Rubrics achieve a score of 54.2% on Qwen3-Coder-30B-A3B and 40.6% on Qwen3-32B, with at least a +3.5 percentage-point gain over the strongest baseline in our comparison set. We further analyze rubric behavior, showing that rubric scores are consistent with ground-truth tests while also flagging issues that tests do not capture. Our ablations show that agentic context gathering is essential for producing codebase-specific, unambiguous criteria. Together, these results suggest that Agentic Rubrics provide an efficient, scalable, and granular verification signal for SWE agents.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
- Rubrics as Rewards: Reinforcement Learning Beyond Verifiable DomainsAnisha Gunjal, Anthony Wang, Elaine Lau, Vaskar Nath 等ICLR 2026 · 被引用 340 次
- SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software EvolutionYuxiang Wei, Olivier Duchenne, Jade Copet, Quentin Carbonneaux 等NeurIPS 2025 · 被引用 291 次
- Checklists Are Better Than Reward Models For Aligning Language ModelsVijay Viswanathan, Yanchao Sun, Xiang Kong, Meng Cao 等NeurIPS 2025 · 被引用 127 次
相关 Paper
- SWE-RM: Execution-free Feedback for Software Engineering AgentsKaShun SHUM, Binyuan Hui, Jiawei Chen, Lei Zhang 等ICLR 2026 · 被引用 24 次
- Scaling Agentic Verifier for Competitive CodingZeyao Ma, Jing Zhang, Xiaokang Zhang, Jiaxi Yang 等ICML 2026 · 被引用 2 次
- CVE-Factory: Scaling Expert-Level Agentic Tasks for Code Security VulnerabilityXianzhen Luo, Jingyuan Zhang, Shiqi Zhou, JinYang Huang 等ICML 2026 · 被引用 3 次
- Training Software Engineering Agents and Verifiers with SWE-GymJiayi Pan, Xingyao Wang, Graham Neubig, Navdeep Jaitly 等ICML 2025
- SWE-rebench V2: Language-Agnostic SWE Task Collection at ScaleIbragim Badertdinov, Maksim Nekrashevich, Anton Shevtsov, Aleksandr GolubevICML 2026 · 被引用 13 次
