Middleware for LLMs: Tools Are Instrumental for Language Agents in Complex Environments
Yu Gu, Yiheng Shu, Hao Yu, Xiao Liu, Yuxiao Dong, Jie Tang, Jayanth Srinivasa, Hugo Latapie, Yu Su
Abstract
The applications of large language models (LLMs) have expanded well beyond the confines of text processing, signaling a new era where LLMs are envisioned as generalist agents capable of operating within complex environments. These environments are often highly expansive, making it impossible for the LLM to process them within its short-term memory. Motivated by recent research on extending the capabilities of LLMs with tools, we seek to investigate the intriguing potential of tools to augment LLMs in handling such complexity by introducing a novel class of tools, termed middleware, to aid in the proactive exploration within these massive environments. Such specialized tools can serve as a middleware layer shielding the LLM from environmental complexity. In two representative complex environmentsknowledge bases (KBs) and databases-we demonstrate the significant potential of augmenting language agents with tools in complex environments. Notably, equipped with the middleware, GPT-4 achieves 2.8× the performance of the best baseline in tasks requiring access to database content and 2.2× in KB tasks. Our findings illuminate the path for advancing language agents in real-world applications. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 17d57ce6-acab-4ba8-9b62-2c9a22afc02fCited by top-tier papers16
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret et al.NeurIPS 2024 · 2,059 citations
- Agent Learning via Early ExperienceKai Zhang, Xiangchao Chen, Bo Liu, Tianci Xue et al.ICML 2026 · 59 citations
- REMem: Reasoning with Episodic Memory in Language AgentYiheng Shu, Padmaja Jonnalagedda, Xiang Gao, Bernal Jimenez Gutierrez et al.ICLR 2026 · 20 citations
- SHARE: An SLM-based Hierarchical Action CorREction Assistant for Text-to-SQLGe Qu, Jinyang Li, Bowen Qin, Xiaolong Li et al.ACL 2025 · 13 citations
- GraphChain: Large Language Models for Large-scale Graph Analysis via Tool ChainingChunyu Wei, Wenji Hu, Xingjia Hao, Xin Wang et al.NeurIPS 2025 · 7 citations
Builds on22
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu et al.NeurIPS 2023 · 5,989 citations
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIsYujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu et al.ICLR 2024 · 1,469 citations
- Teaching Large Language Models to Self-DebugXinyun Chen, Maxwell Lin, Nathanael Schärli, Denny ZhouICLR 2024 · 1,085 citations
- DIN-SQL: Decomposed In-Context Learning of Text-to-SQL with Self-CorrectionMohammadreza Pourreza, Davood RafieiNeurIPS 2023 · 909 citations
- ALFWorld: Aligning Text and Embodied Environments for Interactive LearningMohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk et al.ICLR 2021 · 819 citations
Related papers
- ToolGen: Unified Tool Retrieval and Calling via GenerationRenxi Wang, Xudong Han, Lei Ji, Shu Wang et al.ICLR 2025
- KARL: Reinforcement Learning for LLM Agents on Multi-Turn Knowledge-Intensive Agentic TasksXueqiao Sun, Xiao Liu, Bowen Lv, Hanchen Zhang et al.ACL 2026
- ToolkenGPT: Augmenting Frozen Language Models with Massive Tools via Tool EmbeddingsShibo Hao, Tianyang Liu, Zhen Wang, Zhiting HuNeurIPS 2023 · 315 citations
- CRAFT: Customizing LLMs by Creating and Retrieving from Specialized ToolsetsLifan Yuan, Yangyi Chen, Xingyao Wang, Yi Fung et al.ICLR 2024 · 117 citations
- NaviAgent: Graph‑Driven Bilevel Planning for Scalable Tool OrchestrationYan Jiang, HAO ZHOU, Lizhong Gu, Tianlong Li et al.ICML 2026 · 1 citation
