AdaReason: Progressive Training of Multi-LoRA Adapters for Budget-Adaptive Language Reasoning Models
Jiacheng Wang, Tianle Chen, Pengyu Cheng, Xiaofeng Hou, Jiacheng Liu
Abstract
Large reasoning models (LRMs) have demonstrated remarkable capabilities in solving complex problems through extended chain-of-thought reasoning. However, existing approaches face a fundamental trade-off between computational efficiency and reasoning accuracy. Current methods either lack support for user-specified computational budgets or require maintaining multiple independent models, leading to significant resource overhead. In this paper, we present AdaReason, a unified framework that trains a single base model to support arbitrary user-defined computational budgets through dynamic adapter composition. Our approach introduces three key innovations: (1) a length-adaptive step reward function that stabilizes training across diverse budget constraints, (2) a progressive training strategy that gradually tightens computational bounds while maintaining model performance, and (3) a runtime adapter merging mechanism that dynamically interpolates between different computational preferences. Unlike existing methods that suffer from training instability in large context windows, AdaReason achieves stable convergence through careful reward shaping and progressive constraint tightening. Additionally, we provide a rigorous theoretical analysis, establishing a performance bound for our merged model. Experiments on different reasoning benchmarks demonstrate that AdaReason establishes a new state-of-the-art in the performance-efficiency trade-off and enables flexible runtime budget adaptation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9717bc8a-e1b5-45d4-852a-1b2c5ad034a7Cited by top-tier papers1
Ask how each one uses itBuilds on12
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Training Language Models to Reason EfficientlyDaman Arora, Andrea ZanetteNeurIPS 2025 · 270 citations
- CoT-Valve: Length-Compressible Chain-of-Thought TuningXinyin Ma, Guangnian Wan, Runpeng Yu, Gongfan Fang et al.ACL 2025 · 162 citations
- C3oT: Generating Shorter Chain-of-Thought Without Compromising EffectivenessYu Kang, Xianghui Sun, Liangyu Chen, Wei ZouAAAI 2025 · 162 citations
- Can Language Models Learn to Skip Steps?Tengxiao Liu, Qipeng Guo, Xiangkun Hu, Cheng Jiayang et al.NeurIPS 2024 · 92 citations
Related papers
- Scalable Chain of Thoughts via Elastic ReasoningYuhui Xu, Hanze Dong, Lei Wang, Doyen Sahoo et al.ICLR 2026 · 42 citations
- AdaMix: Adaptive Mixing for Short and Long Reasoning AdaptersHao Luo, Xiao Yan, Xinyan Li, Qiming Zeng et al.ACL 2026
- Pruning Long Chain-of-Thought of Large Reasoning Models via Small-Scale Preference OptimizationBin Hong, Jiayu Liu, Kai Zhang, Jianwen Sun et al.ICLR 2026 · 1 citation
- Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought ReasoningRenos Zabounidis, Aditya Golatkar, Michael Kleinman, Alessandro Achille et al.ICML 2026 · 4 citations
- DeepCompress: A Dual Reward Strategy for Dynamically Exploring and Compressing Reasoning ChainsTian Liang, Wenxiang Jiao, Zhiwei He, Jiahao Xu et al.ICLR 2026 · 10 citations
