Predicting Emergent Tool Use in LLMs Before It Emerges: A Proxy Perspective
Bowen Zhang, Yan Yan, Guang Liu, Xu-Cheng Yin
Abstract
Tool-use capabilities fundamentally transform large language models (LLMs) from passive language generators into active agents with real-world utility, thus drawing intense research focus. However, as a canonical emergent ability characterized by abrupt onset during training, tool-use defies prediction by conventional scaling laws, hindering principled model design and efficient training. In this work, we propose a proxy-task framework to predict emergent tool-use capabilities by measuring early model performance on carefully selected nonemergent tasks. We quantify each proxy task by two properties: alignment, reflecting its correlation with tool-use performance, and consistency, indicating stability across diverse training conditions. These metrics guide a weighted aggregation of proxy signals to predict final tool-use rankings. Theoretically, we formalize how such weighted signals approximate emergent tool use under relaxed assumptions with bounded extrapolation guarantees. Empirically, our approach is validated across training checkpoints, model scales, and data setups. Results demonstrate that a properly weighted ensemble of proxy tasks accurately predicts downstream tooluse ability long before it manifests. Our findings provide new theoretical foundations and practical tools for efficient training and capability planning, advancing understanding of emergent behaviors in LLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 74618164-e578-44dd-a3c8-2fd8bd132684Builds on7
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu et al.NeurIPS 2023 · 5,989 citations
- Gorilla: Large Language Model Connected with Massive APIsShishir G. Patil, Tianjun Zhang, Xin Wang, Joseph E. GonzalezNeurIPS 2024 · 1,715 citations
- Quantifying Memorization Across Neural Language ModelsNicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee et al.ICLR 2023 · 158 citations
- Exploring and Predicting Transferability across NLP TasksTu Vu, Tong Wang, Tsendsuren Munkhdalai, Alessandro Sordoni et al.EMNLP 2020 · 104 citations
Related papers
- U-shaped and Inverted-U Scaling behind Emergent Abilities of Large Language ModelsTung-Yu Wu, Melody LoICLR 2025
- Unveiling Downstream Performance Scaling of LLMs: A Clustering-Based PerspectiveChengyin Xu, Kaiyuan Chen, Xiao Li, Ke Shen et al.ICLR 2026 · 11 citations
- Revisiting the Scaling Properties of Downstream Metrics in Large Language Model TrainingJakub Krajewski, Amitis Shidani, Dan Busbridge, Sam Wiseman et al.ICLR 2026 · 8 citations
- Predicting LLM Reasoning Performance with Small Proxy ModelWoosung Koh, Juyoung Suk, Sungjun Han, Se-Young Yun et al.ICLR 2026 · 5 citations
- Are Emergent Abilities of Large Language Models a Mirage?Rylan Schaeffer, Brando Miranda, Sanmi KoyejoNeurIPS 2023 · 796 citations
