Unveiling Downstream Performance Scaling of LLMs: A Clustering-Based Perspective
Chengyin Xu, Kaiyuan Chen, Xiao Li, Ke Shen, Chenggang Li
Abstract
The escalating scale and cost of Large Language Models (LLMs) training necessitate accurate pre-training prediction of downstream task performance for comprehensive understanding of scaling properties. This is challenged by: 1) the emergence phenomenon, where unpredictable capabilities appearing suddenly at critical model scales; and 2) uneven task difficulty and inconsistent performance scaling patterns, leading to high metric variability. Current prediction methods lack accuracy and reliability. We propose a Clustering-On-Difficulty (COD) framework for downstream performance prediction. The COD framework clusters tasks by their difficulty scaling features, thereby constructing a more stable and predictable task subset that exhibits well-behaved scaling characteristics with the increase of compute budget. We adopt a performance scaling law to predict cluster-wise performance with theoretical support. Predictable subset performance acts as an intermediate predictor for the full evaluation set. We further derive a mapping function to accurately extrapolate the performance of the subset to the full set. Applied to an LLM with 70B parameters, COD achieved a 1.55% average prediction error across eight key LLM benchmarks, thus providing actionable insights for scaling properties and training monitoring during LLM pre-training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 78da587e-687d-4ff0-be01-718ae8da5c76Cited by top-tier papers5
- Learning to Orchestrate Agents in Natural Language with the ConductorStefan Nielsen, Edoardo Cetin, Peter Schwendeman, Qi Sun et al.ICLR 2026 · 22 citations
- Predictable Scale (Part II) - Farseer: A Refined Scaling Law in LLMsHouyi Li, Wenzhen Zheng, Qiufeng Wang, Zhenyu Ding et al.NeurIPS 2025 · 4 citations
- Scaling-Aware Data Selection for End-to-End Autonomous Driving SystemsTolga Dimlioglu, Nadine Chang, Maying Shen, Rafid Mahmood et al.CVPR 2026 · 1 citation
- How2Everything: Mining the Web for How-to Procedures to Evaluate and Improve LLMsYapei Chang, Kyle Lo, Mohit Iyyer, Luca SoldainiICML 2026
- Prescriptive Scaling Reveals the Evolution of Language Model CapabilitiesHanlin Zhang, Jikai Jin, Vasilis Syrgkanis, Sham KakadeICML 2026
Builds on12
- Are Emergent Abilities of Large Language Models a Mirage?Rylan Schaeffer, Brando Miranda, Sanmi KoyejoNeurIPS 2023 · 796 citations
- Scaling Vision TransformersXiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas BeyerCVPR 2022 · 767 citations
- Revisiting Neural Scaling Laws in Language and VisionIbrahim M. Alabdulmohsin, Behnam Neyshabur, Xiaohua ZhaiNeurIPS 2022 · 171 citations
- Scaling Laws with Vocabulary: Larger Models Deserve Larger VocabulariesChaofan Tao, Qian Liu, Longxu Dou, Niklas Muennighoff et al.NeurIPS 2024 · 135 citations
- Understanding Emergent Abilities of Language Models from the Loss PerspectiveZhengxiao Du, Aohan Zeng, Yuxiao Dong, Jie TangNeurIPS 2024 · 113 citations
Related papers
- Revisiting the Scaling Properties of Downstream Metrics in Large Language Model TrainingJakub Krajewski, Amitis Shidani, Dan Busbridge, Sam Wiseman et al.ICLR 2026 · 8 citations
- Predicting Emergent Tool Use in LLMs Before It Emerges: A Proxy PerspectiveBowen Zhang, Yan Yan, Guang Liu, Xu-Cheng YinAAAI 2026
- U-shaped and Inverted-U Scaling behind Emergent Abilities of Large Language ModelsTung-Yu Wu, Melody LoICLR 2025
- Collaborative Performance Prediction for Large Language ModelsQiyuan Zhang, Fuyuan Lyu, Xue Liu, Chen MaEMNLP 2024 · 1 citation
- Sloth: scaling laws for LLM skills to predict multi-benchmark performance across familiesFelipe Maia Polo, Seamus Somerstep, Leshem Choshen, Yuekai Sun et al.NeurIPS 2025 · 28 citations
