Efficient Deweahter Mixture-of-Experts with Uncertainty-Aware Feature-Wise Linear Modulation
Rongyu Zhang, Yulin Luo, Jiaming Liu, Huanrui Yang, Zhen Dong, Denis A. Gudovskiy, Tomoyuki Okuno, Yohei Nakata, Kurt Keutzer, Yuan Du, Shanghang Zhang
Abstract
The Mixture-of-Experts (MoE) approach has demonstrated outstanding scalability in multi-task learning including low-level upstream tasks such as concurrent removal of multiple adverse weather effects. However, the conventional MoE architecture with parallel Feed Forward Network (FFN) experts leads to significant parameter and computational overheads that hinder its efficient deployment. In addition, the naive MoE linear router is suboptimal in assigning task-specific features to multiple experts which limits its further scalability. In this work, we propose an efficient MoE architecture with weight sharing across the experts. Inspired by the idea of linear feature modulation (FM), our architecture implicitly instantiates multiple experts via learnable activation modulations on a single shared expert block. The proposed Feature Modulated Expert (FME) serves as a building block for the novel Mixture-of-Feature-Modulation-Experts (MoFME) architecture, which can scale up the number of experts with low overhead. We further propose an Uncertainty-aware Router (UaR) to assign task-specific features to different FM modules with well-calibrated weights. This enables MoFME to effectively learn diverse expert functions for multiple tasks. The conducted experiments on the multi-deweather task show that our MoFME outperforms the state-of-the-art in the image restoration quality by 0.1-0.2 dB while saving more than 74% of parameters and 20% inference time over the conventional MoE counterpart. Experiments on the downstream segmentation and classification tasks further demonstrate the generalizability of MoFME to real open-world applications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 133f26e2-b886-4450-84a9-cec22d6060f4Cited by top-tier papers11
- SaulLM-54B & SaulLM-141B: Scaling Up Domain Adaptation for the Legal DomainPierre Colombo, Telmo Pessoa Pires, Malik Boudiaf, Rui Melo et al.NeurIPS 2024 · 58 citations
- MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot ManipulationRongyu Zhang, Menghang Dong, Yuan Zhang, Liang Heng et al.AAAI 2026 · 56 citations
- VeCAF: Vision-language Collaborative Active Finetuning with Training Objective AwarenessRongyu Zhang, Zefan Cai, Huanrui Yang, Zidong Liu et al.ACM MM 2024 · 5 citations
- Robust Adverse Weather Removal via Spectral-based Spatial GroupingYuhwan Jeong, Yunseo Yang, Youngho Yoon, Kuk-Jin YoonICCV 2025 · 1 citation
- UAWTrack: Universal 3D Single Object Tracking in Adverse WeatherYuxiang Yang, Hongjie Gu, Yingqi Deng, Zhekang Dong et al.AAAI 2025 · 1 citation
Builds on15
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Restormer: Efficient Transformer for High-Resolution Image RestorationSyed Waqas Zamir, Aditya Arora, Salman Khan, Munawar Hayat et al.CVPR 2022 · 3,348 citations
- Uformer: A General U-Shaped Transformer for Image RestorationZhendong Wang, Xiaodong Cun, Jianmin Bao, Wengang Zhou et al.CVPR 2022 · 1,970 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- FFA-Net: Feature Fusion Attention Network for Single Image DehazingXu Qin, Zhilin Wang, Yuanchao Bai, Xiaodong Xie et al.AAAI 2020 · 1,828 citations
Related papers
- Complexity Experts are Task-Discriminative Learners for Any Image RestorationEduard Zamfir, Zongwei Wu, Nancy Mehta, Yuedong Tan et al.CVPR 2025
- MOERL: When Mixture-Of-Experts Meet Reinforcement Learning for Adverse Weather Image RestorationTao Wang, Peiwen Xia, Bo Li, Peng-Tao Jiang et al.ICCV 2025 · 5 citations
- LiME: Lightweight Mixture of Experts for Efficient Multimodal Multi-task LearningMd Kowsher, Haris Mansoor, Nusrat Prottasha, Ozlem Garibay et al.ICML 2026
- Multi-Task Dense Prediction Fine-Tuning with Mixture of Fine-Grained ExpertsYangyang Xu, Xi Ye, Duo SuACM MM 2025
- MoDE: A Mixture-of-Experts Model with Mutual Distillation among the ExpertsZhitian Xie, Yinger Zhang, Chenyi Zhuang, Qitao Shi et al.AAAI 2024 · 20 citations
