Agent Lumos: Unified and Modular Training for Open-Source Language Agents
Da Yin, Faeze Brahman, Abhilasha Ravichander, Khyathi Raghavi Chandu, Kai-Wei Chang, Yejin Choi, Bill Yuchen Lin
Abstract
Closed-source agents suffer from several issues such as a lack of affordability, transparency, and reproducibility, particularly on complex interactive tasks. This motivates the development of open-source alternatives. We introduce LUMOS, one of the first frameworks for training open-source LLM-based agents. LUMOS features a learnable, unified and modular architecture with a planning module that learns highlevel subgoal generation, and a grounding module trained to translate these into the actions using various tools in the execution module. The design allows for modular upgrades and wider applicability to diverse interactive tasks. To foster generalizable agent learning, we collect large-scale, unified, and high-quality training annotations derived from diverse ground-truth reasoning rationales across various complex interactive tasks. On 9 datasets, LUMOS exhibits several key advantages: (1) LUMOS excels multiple larger open-source agents on the held-out datasets (unused for training) for each task type. LUMOS even surpasses GPT agents on QA and web tasks; (2) LUMOS outperforms opensource agents produced by chain-of-thoughts and unmodularized integrated training; and (3) LUMOS effectively generalizes to unseen tasks, outperforming 33B-scale agents and domainspecific agents. Code and data will be released. closed-source LLMs hinders scientific understand-050 ing of their architectures and effectiveness, and 051 provides limited reproducibility, and controllability 052 over their behavior. We argue that over reliance on 053 closed-source LLM-based agents is not conducive 054 to the growth of research on language agents. 055 In this paper, we propose LUMOS, a gener-056 alizable Language agent framework via Unified, 057 Modular, and Open Source training. LUMOS em-058 ploys a unified and modular architecture broadly 059 applicable to complex interactive tasks: a planning 060 module , a grounding module , and an execu-061 tion module . The planning module learns to 062 decompose diverse complex tasks into a sequence 063 of high-level subgoals. The grounding module is 064 trained to communicate with the planning module 065 ๐ s1: Search flights from honolulu to nyc... โ a1-1: Type([box-id], HNL) โ โ Browser โ a1-2: Type([box-id], JFK) โ โ Browser โ a1-3: Click([button-id]) โ โ Browser ๐ s2: Set a filter to keep โฆ price โค 1300 โ โฆ. โโฆ. ๐ s3: Set a filter to keep premium economy Lumos-OnePass (Lumos-O) Lumos-Iterative (Lumos-I) (๐ก-th iteration) Task desc.: ๐ Prev. subgoals: Prev. actions: Action interfaces: ๐ผ Task desc.: ๐ Prev. results: Prev. subgoals: ๐ Planning โ Grounding โ Execution Task desc: ๐ Action interf.: ๐ผ Task desc.: ๐ All Actions Exe. results All Subgoals ๐ Planning โ Grounding โ Execution Next Actions Exe. result Next Subgoal Multimodal Task (A-OKVQA): The device in her hand is from which country? ๐ s1: Identify the brand of the device โฆ โ a1: VQA(<img>, What is the brand..?) โ e1: LLAVA(...) โ Nintendo Web Task (Mind2Web): Find flights from honolulu to NYC with budget of $1,300 for premium economy. ๐ s2: Answer the country of Nintendo โ a2: QA(context, What's the country โฆ) โ e2: LLM(...) โ Japan 108 web, math, and multimodal tasks. We summarize 109 our contributions and results as follows: 110 General Agent Framework with High-Quality 111 Data. We introduce an open-source agent learn-112 ing framework that trains LLMs with unified data, 113 aimed at unifying complex interactive tasks and 114 enhancing generalization on unseen tasks with new 115 environments and actions. We hope our framework 116 and annotations can facilitate future research in 117 developing open-source language agents. 118 Competitive Performance. LUMOS outper-119 forms a great number of open-source agents on 120 the LUMOS held-out datasets unused in LUMOS 121 training data across the four training task types. 122 LUMOS even surpasses GPT-based agents in web 123 and QA tasks. Specifically, LUMOS shows a 5.0% 124 enhancement over GPT-4 on Mind2Web, and 125 4.1% and 3.5% LLM accuracy 1 improvement on 126 HotpotQA over the ReAct and ReWOO agents 127 fully based on GPT-3.5-turbo, respectively. 128 Cross-Task Generalization. We evaluate LU-129 MOS on two unseen tasks, WebShop (Yao et al., 130 2022a), a text game for online shopping, and 131 InterCode SQL (Yang et al., 2023), an interactive 132 code generation task. LUMOS even surpasses 30B-133 scale agents, especially by nearly 20 reward points 134 on WebShop. LUMOS also delivers a consistent 135 reward improvement over domain-specific agents. 136 This suggests that LUMOS can generalize across 137 tasks, hinting at potential benefits for a wide spec-138 trum of language agent applications. 139 2 LUMOS: A Modular Open-Source 140 LLM-Based Agent Framework 141 We introduce the overall design and two formula-142 tions for developing agents within this framework. 2.1 LUMOS Agent Architecture 144 For various complex interactive tasks, a common 145 solution would include: (1) decomposing the
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aa6bb4cd-db35-43b7-a21d-f6a1b3a397a6Cited by top-tier papers24
- MemGen: Weaving Generative Latent Memory for Self-Evolving AgentsGuibin Zhang, Muxin Fu, Shuicheng YanICLR 2026 ยท 102 citations
- Distilling LLM Agent into Small Models with Retrieval and Code ToolsMinki Kang, Jongwon Jeong, Seanie Lee, Jaewoong Cho et al.NeurIPS 2025 ยท 51 citations
- AMOR: A Recipe for Building Adaptable Modular Knowledge Agents Through Process FeedbackJian Guan, Wei Wu, Zujie Wen, Peng Xu et al.NeurIPS 2024 ยท 35 citations
- Magnet: Multi-turn Tool-use Data Synthesis and Distillation via Graph TranslationFan Yin, Zifeng Wang, I-Hung Hsu, Jun Yan et al.ACL 2025 ยท 23 citations
- AgentRM: Enhancing Agent Generalization with Reward ModelingYu Xia, Jingru Fan, Weize Chen, Siyu Yan et al.ACL 2025 ยท 20 citations
Builds on6
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu et al.NeurIPS 2022 ยท 2,727 citations
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu et al.ACL 2023 ยท 540 citations
- Chameleon: Plug-and-Play Compositional Reasoning with Large Language ModelsPan Lu, Baolin Peng, Hao Cheng, Michel Galley et al.NeurIPS 2023 ยท 515 citations
- Beyond I.I.D.: Three Levels of Generalization for Question Answering on Knowledge BasesYu Gu, Sue Kase, Michelle Vanni, Brian M. Sadler et al.WWW 2021 ยท 304 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 ยท 142 citations
Related papers
- Learning to Contextualize Web Pages for Enhanced Decision Making by LLM AgentsDongjun Lee, Juyong Lee, Kyuyoung Kim, Jihoon Tack et al.ICLR 2025
- AgentRefine: Enhancing Agent Generalization through Refinement TuningDayuan Fu, Keqing He, Yejie Wang, Wentao Hong et al.ICLR 2025
- AutoAct: Automatic Agent Learning from Scratch for QA via Self-PlanningShuofei Qiao, Ningyu Zhang, Runnan Fang, Yujie Luo et al.ACL 2024
- WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement LearningZehan Qi, Xiao Liu, Iat Long Iong, Hanyu Lai et al.ICLR 2025
- AndroidGen: Building an Android Language Agent under Data ScarcityHanyu Lai, Junjie Gao, Xiao Liu, Yifan Xu et al.ACL 2025
