CAMEx: Curvature-aware Merging of Experts
Viet Dung Nguyen, Minh Nguyen Hoang, Luc Q. Nguyen, Rachel S. Y. Teo, Tan Minh Nguyen, Linh Duy Tran
Abstract
Existing methods for merging experts during model training and fine-tuning predominantly rely on Euclidean geometry, which assumes a flat parameter space. This assumption can limit the model's generalization ability, especially during the pre-training phase, where the parameter manifold might exhibit more complex curvature. Curvature-aware merging methods typically require additional information and computational resources to approximate the Fisher Information Matrix, adding memory overhead. In this paper, we introduce CAMEx (Curvature-Aware Merging of Experts), a novel expert merging protocol that incorporates natural gradients to account for the non-Euclidean curvature of the parameter manifold. By leveraging natural gradients, CAMEx adapts more effectively to the structure of the parameter space, improving alignment between model updates and the manifold's geometry. This approach enhances both pre-training and fine-tuning, resulting in better optimization trajectories and improved generalization without the substantial memory overhead typically associated with curvatureaware methods. Our contributions are two-fold: (1) CAMEx significantly outperforms traditional Euclidean-based expert merging techniques across various natural language processing tasks, leading to enhanced performance during pretraining and fine-tuning; (2) we introduce a dynamic merging architecture that optimizes resource utilization, achieving high performance while reducing computational costs, facilitating efficient scaling of large language models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Expert Merging in Sparse Mixture of Experts with Nash BargainingDung Viet Nguyen, Anh Nguyen Thi, Minh Hoang Nguyen, Luc Nguyen et al.ICLR 2026 · 3 citations
- MoLEx: Mixture of Layer Experts for Fine-tuning with Sparse UpcyclingRachel S. Y. Teo, Tan Minh NguyenICLR 2025
- Tight Clusters Make Specialized ExpertsStefan K. Nielsen, Rachel S. Y. Teo, Laziz U. Abdullaev, Tan Minh NguyenICLR 2025
Builds on23
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath et al.ICCV 2021 · 2,294 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- Mixture-of-Experts with Expert Choice RoutingYanqi Zhou, Tao Lei, Hanxiao Liu, Nan Du et al.NeurIPS 2022 · 933 citations
Related papers
- MergOPT: A Merge-Aware Optimizer for Robust Model MergingEnneng Yang, Qun Yang, Peng Wang, Anke Tang et al.ICLR 2026
- When Model Merging Breaks Routing: Training-Free Calibration for MoECanbin Huang, Tianyuan Shi, Xiaojun Quan, Jingang Wang et al.ICML 2026
- Extra-Merge: Tracing the Rank-1 Subspace of Model Merging in Language Model Pre-TrainingWenJie Zhou, Bohan Wang, Hongtao Zhang, Chenxi Jia et al.ICML 2026
- A Data-Efficient Path to Multilingual LLMs: Language Expansion via Post-training PARAMΔ Integration into Upcycled MoEHao Zhou, Tianhao Li, Zhijun Wang, Shuaijie She et al.ACL 2026
- Expert Merging: Model Merging with Unsupervised Expert Alignment and Importance-Guided Layer ChunkingDengming Zhang, Xiaowen Ma, Zhenliang Ni, Zhenkai Wu et al.ICLR 2026 · 6 citations
