Predicting Emergent Tool Use in LLMs Before It Emerges: A Proxy Perspective
Bowen Zhang, Yan Yan, Guang Liu, Xu-Cheng Yin
摘要
Tool-use capabilities fundamentally transform large language models (LLMs) from passive language generators into active agents with real-world utility, thus drawing intense research focus. However, as a canonical emergent ability characterized by abrupt onset during training, tool-use defies prediction by conventional scaling laws, hindering principled model design and efficient training. In this work, we propose a proxy-task framework to predict emergent tool-use capabilities by measuring early model performance on carefully selected nonemergent tasks. We quantify each proxy task by two properties: alignment, reflecting its correlation with tool-use performance, and consistency, indicating stability across diverse training conditions. These metrics guide a weighted aggregation of proxy signals to predict final tool-use rankings. Theoretically, we formalize how such weighted signals approximate emergent tool use under relaxed assumptions with bounded extrapolation guarantees. Empirically, our approach is validated across training checkpoints, model scales, and data setups. Results demonstrate that a properly weighted ensemble of proxy tasks accurately predicts downstream tooluse ability long before it manifests. Our findings provide new theoretical foundations and practical tools for efficient training and capability planning, advancing understanding of emergent behaviors in LLMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu 等NeurIPS 2023 · 被引用 5,989 次
- Gorilla: Large Language Model Connected with Massive APIsShishir G. Patil, Tianjun Zhang, Xin Wang, Joseph E. GonzalezNeurIPS 2024 · 被引用 1,715 次
- Quantifying Memorization Across Neural Language ModelsNicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee 等ICLR 2023 · 被引用 158 次
- Exploring and Predicting Transferability across NLP TasksTu Vu, Tong Wang, Tsendsuren Munkhdalai, Alessandro Sordoni 等EMNLP 2020 · 被引用 104 次
相关 Paper
- U-shaped and Inverted-U Scaling behind Emergent Abilities of Large Language ModelsTung-Yu Wu, Melody LoICLR 2025
- Unveiling Downstream Performance Scaling of LLMs: A Clustering-Based PerspectiveChengyin Xu, Kaiyuan Chen, Xiao Li, Ke Shen 等ICLR 2026 · 被引用 11 次
- Revisiting the Scaling Properties of Downstream Metrics in Large Language Model TrainingJakub Krajewski, Amitis Shidani, Dan Busbridge, Sam Wiseman 等ICLR 2026 · 被引用 8 次
- Predicting LLM Reasoning Performance with Small Proxy ModelWoosung Koh, Juyoung Suk, Sungjun Han, Se-Young Yun 等ICLR 2026 · 被引用 5 次
- Are Emergent Abilities of Large Language Models a Mirage?Rylan Schaeffer, Brando Miranda, Sanmi KoyejoNeurIPS 2023 · 被引用 796 次
