OctoTools: A Multi-Agent Framework with Extensible Tools for Complex Reasoning
Pan Lu, Bowen Chen, Sheng Liu, Rahul Thapa, Joseph Boen, James Zou
摘要
Solving complex reasoning tasks may involve visual understanding, domain knowledge retrieval, numerical calculation, and multi-step reasoning.Existing methods augment large language models (LLMs) with external tools but are restricted to specialized domains, limited tool types, or require additional training data.In this paper, we introduce OctoTools, a training-free, user-friendly, and easily extensible multi-agent framework designed to tackle complex reasoning across diverse domains.Oc-toTools introduces standardized tool cards to encapsulate tool functionality, a planner for both high-level and low-level planning, and an executor to carry out tool usage.We validate OctoTools' generality across 16 diverse tasks (including MathVista, MMLU-Pro, MedQA, and GAIA-Text), achieving substantial average accuracy gains of 9.3% over GPT-4o.Furthermore, OctoTools also outperforms AutoGen, GPT-Functions, and LangChain by up to 10.6% when given the same set of tools.Through comprehensive analysi, ablations, and robustness tests with compact backbones and noisy tool environments, OctoTools demonstrates advantages in task planning, effective tool usage, and multi-step problem solving.Q: How many baseballs are there?The image shows four blue buckets, each containing five baseballs.Therefore, there are a total of 20 baseballs. Solution Summarizer
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu 等NeurIPS 2023 · 被引用 5,989 次
- GAIA: a benchmark for General AI AssistantsGrégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun 等ICLR 2024 · 被引用 716 次
相关 Paper
- AdaReasoner: Dynamic Tool Orchestration for Iterative Visual ReasoningMingyang Song, Haoyu Sun, Jiawei Gu, Linjie Li 等ICLR 2026 · 被引用 7 次
- CRAFT: Customizing LLMs by Creating and Retrieving from Specialized ToolsetsLifan Yuan, Yangyi Chen, Xingyao Wang, Yi Fung 等ICLR 2024 · 被引用 117 次
- PlanningArena: A Modular Benchmark for Multidimensional Evaluation of Planning and Tool LearningZihan Zheng, Tianle Cui, Chuwen Xie, Jiahui Pan 等ACL 2025 · 被引用 3 次
- TUMIX: Multi-Agent Test-Time Scaling with Tool-Use MixtureYongchao Chen, Jiefeng Chen, Rui Meng, Ji Yin 等ICLR 2026 · 被引用 13 次
- THOR: Tool-Integrated Hierarchical Optimization via RL for Mathematical ReasoningQikai Chang, Zhenrong Zhang, Pengfei Hu, Jun Du 等ICLR 2026 · 被引用 8 次
