Compute Where it Counts: Self Optimizing Language Models
Yash Akhauri, Mohamed Abdelfattah
摘要
Efficient LLM inference research has largely focused on reducing the cost of each decoding step (e.g., using quantization, pruning, or sparse attention), typically applying a uniform computation budget to every generated token. In practice, token difficulty varies widely, so static compression can over-compute on easy steps and under-compute on hard ones. We study dynamic budget allocation for autoregressive decoding: learning how much computation to spend per token from within a single model. Self-Optimizing Language Models (SOL) pair a frozen LLM with a lightweight policy network that reads the LLM hidden state and selects a discrete efficiency action at each decode step. Actions can jointly control (i) token-level attention sparsity, (ii) structured activation pruning in the MLP, and (iii) activation quantization bit-width, while leaving the base model weights unchanged. We train the policy with group-relative policy optimization on teacher-forced episodes: the token sequence is fixed, while we sample multiple compute schedules (i.e., “counterfactual” schedules that vary only the efficiency actions for the same token path) and compare their likelihoods under the same supervision. Our reward trades off language‑model quality against soft penalties that encourage episode‑average budget usage to match a requested target. Across model variants and compute regimes, SOL improves quality at matched budget over static allocation and strong random schedule search, offering a complementary axis for inference‑efficiency optimization. SOL discovers a better quality-efficiency pareto-front across all our experiments and improves MMLU accuracy by up to 7.3% over uniform budget allocation strategies. https://github.com/akhauriyash/SOL
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
相关 Paper
- Scheduling Thoughts: Learning the Order of Thought in Diffusion Language ModelsJiawei Xu, Minghui Liu, Aakriti Agrawal, Yifan Chen 等ICML 2026 · 被引用 1 次
- Learning How Hard to Think: Input-Adaptive Allocation of LM ComputationMehul Damani, Idan Shenfeld, Andi Peng, Andreea Bobu 等ICLR 2025
- Sparsity Forcing: Reinforcing Token Sparsity of MLLMsFeng Chen, Yefei He, Lequan Lin, Jing Liu 等ICLR 2026 · 被引用 3 次
- Let LLM Tell What to Prune and How Much to PruneMingzhe Yang, Sihao Lin, Changlin Li, Xiaojun ChangICML 2025
- DISC: Dynamic Decomposition Improves LLM Inference ScalingJonathan Light, Wei Cheng, Benjamin Rivière, Yue Wu 等NeurIPS 2025 · 被引用 12 次
