RAAS: LLM Agentic System Architecture Search with GRPO
Jiayi Yang, Guancheng Wan, Man Zhang, Mang Ye
Abstract
Large Language Model (LLM) agentic systems solve complex tasks through coordinated workflows, but designing them remains labor-intensive. The Agentic Supernet paradigm automates this by optimizing a probabilistic architecture space, yet suffers from critical evaluation instabilities: absolute performance scores entangle architectural merit with query difficulty, while single-execution protocols capture execution randomness rather than true capability. These instabilities lead to unreliable search dynamics where simple queries inflate weak designs and challenging queries suppress strong ones. We introduce RAAS (Robust Architecture Adaptive Search), which establishes more stable and fair evaluation through two synergistic mechanisms. Contextual Architecture Orchestration (CAO) disentangles quality from task difficulty by evaluating cohorts of candidate architectures on identical queries, deriving context-aware merit signals through peer-group comparison. Multi-Trial Assessment Synthesis (MTAS) reduces execution variance by aggregating performance across multiple independent trials, producing statistically robust capability estimates. Together, these mechanisms provide more reliable signals for architecture discovery. Experiments on six benchmarks spanning mathematical reasoning, code generation, and one multi-step tool-use benchmark show that RAAS consistently improves over strong baselines, improving HumanEval pass@1 from 92.23% to 96.31% and MATH accuracy from 52.08% to 60.87%, while maintaining practical efficiency. These results suggest that robust evaluation is a useful ingredient for agentic architecture search in the studied settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5032c929-e86a-41ee-ad45-5dd4e17cb89eBuilds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging FaceYongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li et al.NeurIPS 2023 · 1,778 citations
Related papers
- Multi-agent Architecture Search via Agentic SupernetGuibin Zhang, Luyang Niu, Junfeng Fang, Kun Wang et al.ICML 2025
- AutoRAS: Learning Robust Agentic Systems with Primitive RepresentationsYang Yue, Xuancheng Zhu, YuYang Ma, Guoshun Nan et al.ICML 2026
- Difficulty-Aware Agentic Orchestration for Query-Specific Multi-Agent WorkflowsJinwei Su, Qizhen Lan, Yinghui Xia, Lifan Sun et al.WWW 2026 · 8 citations
- SwarmAgentic: Towards Fully Automated Agentic System Generation via Swarm IntelligenceYao Zhang, Chenyang Lin, Shijie Tang, Haokun Chen et al.EMNLP 2025 · 2 citations
- EvoMAS: Evolutionary Generation of Multi-Agent SystemsYuntong Hu, Yuting Zhang, Matthew Trager, Yi Zhang et al.ICML 2026
