TroVE: Inducing Verifiable and Efficient Toolboxes for Solving Programmatic Tasks
Zhiruo Wang, Graham Neubig, Daniel Fried
摘要
Language models (LMs) can solve tasks such as answering questions about tables or images by writing programs. However, using primitive functions often leads to verbose and error-prone programs, and higher-level functions require expert design. To enable better solutions without human labor, we ask code LMs to curate reusable high-level functions, and use them to write solutions. We present TROVE, a training-free method of inducing a verifiable and efficient toolbox of functions, by generating via using, growing, and periodically trimming the toolbox. On 11 datasets from math, table question answering, and image reasoning tasks, TROVE consistently yields simpler solutions with higher accuracy than baselines using CODELLAMA and previous methods using GPT, while using 79-98% smaller toolboxes. TROVE further enables 31% faster and 13% more accurate human verification than baselines. With the same pipeline, it creates diverse functions for varied tasks and datasets, providing insights into their individual characteristics. Code and data are available at https://github.com/zorazrw/trove .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret 等NeurIPS 2024 · 被引用 2,059 次
- SE-Agent: Self-Evolution Trajectory Optimization in Multi-Step Reasoning with LLM-Based AgentsYifu Guo, Jiaye Lin, Huacan Wang, Yuzhen Han 等NeurIPS 2025 · 被引用 73 次
- RepoMaster: Autonomous Exploration and Understanding of GitHub Repositories for Complex Task SolvingHuacan Wang, Ziyi Ni, Shuo Zhang, Shuo Lu 等NeurIPS 2025 · 被引用 27 次
- ReGAL: Refactoring Programs to Discover Generalizable AbstractionsElias Stengel-Eskin, Archiki Prasad, Mohit BansalICML 2024 · 被引用 23 次
- SciAgent: Tool-augmented Language Models for Scientific ReasoningYubo Ma, Zhibin Gou, Junheng Hao, Ruochen Xu 等EMNLP 2024 · 被引用 13 次
它引用的顶会 Paper13
- ViperGPT: Visual Inference via Python Execution for ReasoningDídac Surís, Sachit Menon, Carl VondrickICCV 2023 · 被引用 732 次
- PAL: Program-aided Language ModelsLuyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon 等ICML 2023 · 被引用 700 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
- Large Language Models as Tool MakersTianle Cai, Xuezhi Wang, Tengyu Ma, Xinyun Chen 等ICLR 2024 · 被引用 283 次
- LEGO-Prover: Neural Theorem Proving with Growing LibrariesHaiming Wang, Huajian Xin, Chuanyang Zheng, Zhengying Liu 等ICLR 2024 · 被引用 125 次
相关 Paper
- CRAFT: Customizing LLMs by Creating and Retrieving from Specialized ToolsetsLifan Yuan, Yangyi Chen, Xingyao Wang, Yi Fung 等ICLR 2024 · 被引用 117 次
- GPT4Tools: Teaching Large Language Model to Use Tools via Self-instructionRui Yang, Lin Song, Yanwei Li, Sijie Zhao 等NeurIPS 2023 · 被引用 340 次
- Language Models of Code are Few-Shot Commonsense LearnersAman Madaan, Shuyan Zhou, Uri Alon, Yiming Yang 等EMNLP 2022 · 被引用 103 次
- SolverLLM: Leveraging Test-Time Scaling for Optimization Problem via LLM-Guided SearchDong Li, Xujiang Zhao, Linlin Yu, Yanchi Liu 等NeurIPS 2025 · 被引用 15 次
- API Pack: A Massive Multi-Programming Language Dataset for API Call GenerationZhen Guo, Adriana Meza Soria, Wei Sun, Yikang Shen 等ICLR 2025
