Activation-Guided Consensus Merging for Large Language Models
Yuxuan Yao, Shuqi Liu, Zehua Liu, Qintong Li, Mingyang Liu, Xiongwei Han, Zhijiang Guo, Han Wu, Linqi Song
摘要
Recent research has increasingly focused on reconciling the reasoning capabilities of System 2 with the efficiency of System 1. While existing training-based and prompt-based approaches face significant challenges in terms of efficiency and stability, model merging emerges as a promising strategy to integrate the diverse capabilities of different Large Language Models (LLMs) into a unified model. However, conventional model merging methods often assume uniform importance across layers, overlooking the functional heterogeneity inherent in neural components. To address this limitation, we propose Activation-Guided Consensus Merging (ACM), a plug-and-play merging framework that determines layer-specific merging coefficients based on mutual information between activations of pre-trained and fine-tuned models. ACM effectively preserves task-specific capabilities without requiring gradient computations or additional training. Extensive experiments on Long-to-Short (L2S) and general merging tasks demonstrate that ACM consistently outperforms all baseline methods. For instance, in the case of Qwen-7B models, TIES-Merging equipped with ACM achieves a 55.3% reduction in response length while simultaneously improving reasoning accuracy by 1.3 points.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- RAIN-Merging: A Gradient-Free Method to Enhance Instruction Following in Large Reasoning Models with Preserved Thinking FormatZhehao Huang, Yuhang Liu, Baijiong Lin, Yixin Lou 等ICLR 2026 · 被引用 7 次
- Too Long, Do Re-weighting for Efficient LLM Reasoning CompressionZhong-Zhi Li, Xiao Liang, Zihao Tang, Lei Ji 等ACL 2026 · 被引用 5 次
- Label-Free Cross-Task LoRA Merging with Null-Space CompressionWonyoung Lee, Wooseong Jeong, Kuk-Jin YoonCVPR 2026 · 被引用 3 次
- Sparsity Curse: Understanding RLVR Model Parameter Space from Model MergingChenrui Wu, Zexi Li, Jiajun Bu, Jiangchuan Liu 等KDD 2026 · 被引用 2 次
- EvoGM: Learning to Merge LLMs via Evolutionary Generative OptimizationTao Jiang, Xinmeng Yu, Chenhao Yi, Yiling Wu 等ICML 2026 · 被引用 1 次
它引用的顶会 Paper16
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer 等NeurIPS 2022 · 被引用 2,039 次
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs 等ICML 2022 · 被引用 1,464 次
- TIES-Merging: Resolving Interference When Merging ModelsPrateek Yadav, Derek Tam, Leshem Choshen, Colin A. Raffel 等NeurIPS 2023 · 被引用 999 次
- Language Models are Super Mario: Absorbing Abilities from Homologous Models as a Free LunchLe Yu, Bowen Yu, Haiyang Yu, Fei Huang 等ICML 2024 · 被引用 605 次
相关 Paper
- Outlier Matters: Efficient Long-to-Short Reasoning via Outlier-Guided Model MergingQiyuan Zhu, Dezhi Li, Lujun Li, Xiaoyu Qin 等AAAI 2026
- Beyond Layer-Wise Merging: Chain-of-Merging for Vision-Language ModelsXinyu Zhang, Yuxuan Dong, Lingling Zhang, Chengyou Jia 等CVPR 2026
- RCP-Merging: Merging Long Chain-of-Thought Models with Domain-Specific Models by Considering Reasoning Capability as PriorJunyao Yang, Jianwei Wang, Huiping Zhuang, Cen Chen 等AAAI 2026 · 被引用 1 次
- Activation-Informed Merging of Large Language ModelsAmin Heyrani Nobari, Kaveh Alimohammadi, Ali ArjomandBigdeli, Akash Srivastava 等NeurIPS 2025 · 被引用 25 次
- Layer Swapping for Zero-Shot Cross-Lingual Transfer in Large Language ModelsLucas Bandarkar, Benjamin Muller, Pritish Yuvraj, Rui Hou 等ICLR 2025
