Efficient Deweahter Mixture-of-Experts with Uncertainty-Aware Feature-Wise Linear Modulation
Rongyu Zhang, Yulin Luo, Jiaming Liu, Huanrui Yang, Zhen Dong, Denis A. Gudovskiy, Tomoyuki Okuno, Yohei Nakata, Kurt Keutzer, Yuan Du, Shanghang Zhang
摘要
The Mixture-of-Experts (MoE) approach has demonstrated outstanding scalability in multi-task learning including low-level upstream tasks such as concurrent removal of multiple adverse weather effects. However, the conventional MoE architecture with parallel Feed Forward Network (FFN) experts leads to significant parameter and computational overheads that hinder its efficient deployment. In addition, the naive MoE linear router is suboptimal in assigning task-specific features to multiple experts which limits its further scalability. In this work, we propose an efficient MoE architecture with weight sharing across the experts. Inspired by the idea of linear feature modulation (FM), our architecture implicitly instantiates multiple experts via learnable activation modulations on a single shared expert block. The proposed Feature Modulated Expert (FME) serves as a building block for the novel Mixture-of-Feature-Modulation-Experts (MoFME) architecture, which can scale up the number of experts with low overhead. We further propose an Uncertainty-aware Router (UaR) to assign task-specific features to different FM modules with well-calibrated weights. This enables MoFME to effectively learn diverse expert functions for multiple tasks. The conducted experiments on the multi-deweather task show that our MoFME outperforms the state-of-the-art in the image restoration quality by 0.1-0.2 dB while saving more than 74% of parameters and 20% inference time over the conventional MoE counterpart. Experiments on the downstream segmentation and classification tasks further demonstrate the generalizability of MoFME to real open-world applications.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- SaulLM-54B & SaulLM-141B: Scaling Up Domain Adaptation for the Legal DomainPierre Colombo, Telmo Pessoa Pires, Malik Boudiaf, Rui Melo 等NeurIPS 2024 · 被引用 58 次
- MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot ManipulationRongyu Zhang, Menghang Dong, Yuan Zhang, Liang Heng 等AAAI 2026 · 被引用 56 次
- VeCAF: Vision-language Collaborative Active Finetuning with Training Objective AwarenessRongyu Zhang, Zefan Cai, Huanrui Yang, Zidong Liu 等ACM MM 2024 · 被引用 5 次
- Robust Adverse Weather Removal via Spectral-based Spatial GroupingYuhwan Jeong, Yunseo Yang, Youngho Yoon, Kuk-Jin YoonICCV 2025 · 被引用 1 次
- UAWTrack: Universal 3D Single Object Tracking in Adverse WeatherYuxiang Yang, Hongjie Gu, Yingqi Deng, Zhekang Dong 等AAAI 2025 · 被引用 1 次
它引用的顶会 Paper15
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Restormer: Efficient Transformer for High-Resolution Image RestorationSyed Waqas Zamir, Aditya Arora, Salman Khan, Munawar Hayat 等CVPR 2022 · 被引用 3,348 次
- Uformer: A General U-Shaped Transformer for Image RestorationZhendong Wang, Xiaodong Cun, Jianmin Bao, Wengang Zhou 等CVPR 2022 · 被引用 1,970 次
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen 等ICLR 2021 · 被引用 1,954 次
- FFA-Net: Feature Fusion Attention Network for Single Image DehazingXu Qin, Zhilin Wang, Yuanchao Bai, Xiaodong Xie 等AAAI 2020 · 被引用 1,828 次
相关 Paper
- Complexity Experts are Task-Discriminative Learners for Any Image RestorationEduard Zamfir, Zongwei Wu, Nancy Mehta, Yuedong Tan 等CVPR 2025
- MOERL: When Mixture-Of-Experts Meet Reinforcement Learning for Adverse Weather Image RestorationTao Wang, Peiwen Xia, Bo Li, Peng-Tao Jiang 等ICCV 2025 · 被引用 5 次
- LiME: Lightweight Mixture of Experts for Efficient Multimodal Multi-task LearningMd Kowsher, Haris Mansoor, Nusrat Prottasha, Ozlem Garibay 等ICML 2026
- Multi-Task Dense Prediction Fine-Tuning with Mixture of Fine-Grained ExpertsYangyang Xu, Xi Ye, Duo SuACM MM 2025
- MoDE: A Mixture-of-Experts Model with Mutual Distillation among the ExpertsZhitian Xie, Yinger Zhang, Chenyi Zhuang, Qitao Shi 等AAAI 2024 · 被引用 20 次
