Parallel Scaling Law for Language Models
Mouxiang Chen, Binyuan Hui, Zeyu Cui, Jiaxi Yang, Dayiheng Liu, Jianling Sun, Junyang Lin, Zhongxin Liu
摘要
It is commonly believed that scaling language models should commit a significant space or time cost, by increasing the parameters (parameter scaling) or output tokens (inference-time scaling). We introduce another and more inference-efficient scaling paradigm: increasing the model's parallel computation during both training and inference time. We apply P diverse and learnable transformations to the input, execute forward passes of the model in parallel, and dynamically aggregate the P outputs. This method, namely parallel scaling (PARSCALE), scales parallel computation by reusing existing parameters and can be applied to any model structure, optimization procedure, data, or task. We theoretically propose a new scaling law and validate it through large-scale pre-training, which shows that a model with P parallel streams is similar to scaling the parameters by O(log P ) while showing superior inference efficiency. For example, PARSCALE can use up to 22× less memory increase and 6× less latency increase compared to parameter scaling that achieves the same performance improvement. It can also recycle an off-the-shelf pre-trained model into a parallelly scaled one by post-training on a small amount of tokens, further reducing the training budget. The new scaling law we discovered potentially facilitates the deployment of more powerful models in low-resource scenarios, and provides an alternative perspective for the role of computation in machine learning. Our code and 67 trained model checkpoints are publicly available at https://github.com/QwenLM/ParScale and https://huggingface.co/ParScale.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Hogwild! Inference: Parallel LLM Generation via Concurrent AttentionGleb Rodionov, Roman Garipov, Alina Shutova, George Yakushev 等NeurIPS 2025 · 被引用 35 次
- Kimi-Dev: Agentless Training as Skill Prior for SWE-agentsZonghan Yang, Shengjie Wang, Kelin Fu, Wenyang He 等ICLR 2026 · 被引用 34 次
- Can Language Models Discover Scaling Laws?Haowei Lin, Haotian Ye, Wenzheng Feng, Quzhe Huang 等ICLR 2026 · 被引用 11 次
- ATTS: Asynchronous Test-Time Scaling via Conformal PredictionJing Xiong, Qiujiang Chen, Fanghua Ye, Zhongwei Wan 等ICLR 2026 · 被引用 8 次
- Generalized Parallel Scaling with Interdependent GenerationsHarry Dong, David Brandfonbrener, Eryk Helenowski, Yun He 等ICLR 2026 · 被引用 7 次
它引用的顶会 Paper39
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
相关 Paper
- Thoughtbubbles: an Unsupervised Method for Parallel Thinking in Latent SpaceHoujun Liu, Shikhar Murty, Christopher Manning, Róbert CsordásICML 2026 · 被引用 3 次
- Scaling Laws for PrecisionTanishq Kumar, Zachary Ankner, Benjamin Frederick Spector, Blake Bordelon 等ICLR 2025
- LESA: Learnable LLM Layer Scaling-UpYifei Yang, Zouying Cao, Xinbei Ma, Yao Yao 等ACL 2025 · 被引用 6 次
- Pre-training under infinite computeKonwoo Kim, Suhas Kotha, Percy Liang, Tatsunori HashimotoICLR 2026 · 被引用 25 次
- Scaling Retrieval-Based Language Models with a Trillion-Token DatastoreRulin Shao, Jacqueline He, Akari Asai, Weijia Shi 等NeurIPS 2024 · 被引用 76 次
