OPIC: Enhancing Language Model Merging via Optimizing In-Context Capability
Jie He, Weidong Bao, Chao Chen, Zhengyi Zhong, Shuai Zhang, Ji Wang
摘要
Task-vector-based model merging enables lowcost, training-free multi-task learning for large language models, but suffers from performance degradation compared to individually fine-tuned models. Prior mitigation strategies largely rely on validation data for costly hyperparameter tuning, limiting both interpretability and practicality. We therefore propose OPIC, an evolutionary optimization-based model merging framework. Our preliminary experiments reveal that the degradation of in-context learning (ICL) capabilities exhibits a strong correlation with performance deterioration. Motivated by this insight, we formulate model merging as an optimization problem with ICL preservation as the objective. OPIC introduces a hierarchical refinement operators and optimizes it using self-generated data, effectively eliminating the reliance on external validation sets. Experimental results demonstrate that OPIC achieves an average performance retention of 80.73%, outperforming SOTA methods and improving by up to 11.1% over recent validation-free approaches. In addition, OPIC is compatible with existing merging pipelines, offering a new alternative solution for deploying without validation dependencies. Code is available at: https://github.com/illusion-hj/OPIC
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach 等ICLR 2022 · 被引用 1,976 次
- TIES-Merging: Resolving Interference When Merging ModelsPrateek Yadav, Derek Tam, Leshem Choshen, Colin A. Raffel 等NeurIPS 2023 · 被引用 999 次
- Merging Models with Fisher-Weighted AveragingMichael Matena, Colin RaffelNeurIPS 2022 · 被引用 741 次
相关 Paper
- Unraveling LoRA Interference: Orthogonal Subspaces for Robust Model MergingHaobo Zhang, Jiayu ZhouACL 2025
- Multi-Modality Expansion and Retention for LLMs through Parameter Merging and DecouplingJunlin Li, Guodong Du, Jing Li, Sim Kuan Goh 等ACL 2025
- Expert Merging: Model Merging with Unsupervised Expert Alignment and Importance-Guided Layer ChunkingDengming Zhang, Xiaowen Ma, Zhenliang Ni, Zhenkai Wu 等ICLR 2026 · 被引用 6 次
- Mimic In-Context Learning for Multimodal TasksYuchu Jiang, Jiale Fu, Chenduo Hao, Xinting Hu 等CVPR 2025
- Merging on the Fly Without Retraining: A Sequential Approach to Scalable Continual Model MergingAnke Tang, Enneng Yang, Li Shen, Yong Luo 等NeurIPS 2025 · 被引用 8 次
