Optimas: Optimizing Compound AI Systems with Globally Aligned Local Rewards
Shirley Wu, Parth Sarthi, Shiyu Zhao, Aaron Lee, Herumb Shandilya, Adrian Mladenic Grobelnik, Nurendra Choudhary, Edward W Huang, Karthik Subbian, Linjun Zhang, Diyi Yang, James Zou, Jure Leskovec
摘要
Compound AI systems integrating multiple components, such as Large Language Models, specialized tools, and traditional machine learning models, are increasingly deployed to solve complex real-world tasks. However, optimizing compound systems remains challenging due to their non-differentiable structures and diverse configuration types across components, including prompts, hyperparameters, and model parameters. To address this challenge, we propose Optimas, a unified framework for effective optimization of compound systems. The core idea of Optimas is to maintain one Local Reward Function (LRF) per component, each satisfying a local–global alignment property, i.e., each component’s local reward correlates with the global system performance. In each iteration, Optimas efficiently adapts the LRFs to maintain this property while simultaneously maximizing each component’s local reward. This approach enables independent updates of heterogeneous configurations using the designated optimization method, while ensuring that local improvements consistently lead to performance gains. We present extensive evaluations across five real-world compound systems to demonstrate that Optimas outperforms strong baselines by an average improvement of 11.92%, offering a general and effective approach for improving compound systems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- GEPA: Reflective Prompt Evolution Can Outperform Reinforcement LearningLakshya A. Agrawal, Shangyin Tan, Dilara Soylu, Noah Ziems 等ICLR 2026 · 被引用 466 次
- Murakkab: Resource-Efficient Agentic Workflow Orchestration in Cloud PlatformsGohar Irfan Chaudhry, Esha Choukse, Haoran Qiu, Íñigo Goiri 等OSDI 2026 · 被引用 26 次
- EvoRoute: Experience-Driven Self-Routing LLM Agent SystemsGuibin Zhang, Haiyang Yu, Kaiming Yang, Bingli Wu 等ACL 2026 · 被引用 6 次
- Aligning Compound AI Systems via System-level DPOXiangwen Wang, Yibo Jacky Zhang, Zhoujie Ding, Katherine Tsai 等NeurIPS 2025 · 被引用 4 次
它引用的顶会 Paper25
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
相关 Paper
- Compound AI Systems Optimization: A Survey of Methods, Challenges, and Future DirectionsYu-Ang Lee, Guan-Ting Yi, Mei-Yi Liu, Jui-Chao Lu 等EMNLP 2025
- Towards AutoAI: Optimizing a Machine Learning System with Black-box and Differentiable ComponentsZhiliang Chen, Chuan-Sheng Foo, Bryan Kian Hsiang LowICML 2024 · 被引用 10 次
- Optimas: An Intelligent Analytics-Informed Generative AI Framework for Performance OptimizationMohammad Zaeed, Tanzima Z. Islam, Vladimir IndicKDD 2026
- metaTextGrad: Automatically optimizing language model optimizersGuowei Xu, Mert Yüksekgönül, Carlos Guestrin, James Y. ZouNeurIPS 2025 · 被引用 5 次
- SCOPE: Cost-Efficient Model Selection for Compound AI Systems under Quality ConstraintsYiqian Huang, Shiqi Zhang, Tianyuan Jin, Xiaokui XiaoKDD 2026
