DAMEX: Dataset-aware Mixture-of-Experts for visual understanding of mixture-of-datasets
Yash Jain, Harkirat S. Behl, Zsolt Kira, Vibhav Vineet
摘要
Construction of a universal detector poses a crucial question: How can we most effectively train a model on a large mixture of datasets? The answer lies in learning dataset-specific features and ensembling their knowledge but do all this in a single model. Previous methods achieve this by having separate detection heads on a common backbone but that results in a significant increase in parameters. In this work, we present Mixture-of-Experts as a solution, highlighting that MoEs are much more than a scalability tool. We propose Dataset-Aware Mixture-of-Experts, DAMEX where we train the experts to become an 'expert' of a dataset by learning to route each dataset tokens to its mapped expert. Experiments on Universal Object-Detection Benchmark show that we outperform the existing state-of-the-art by average +10.2 AP score and improve over our non-MoE baseline by average +2.0 AP score. We also observe consistent gains while mixing datasets with (1) limited availability, (2) disparate domains and (3) divergent label sets. Further, we qualitatively show that DAMEX is robust against expert representation collapse. Code is available at https://github.com/jinga-lala/DAMEX .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- SM3Det: A Unified Model for Multi-Modal Remote Sensing Object DetectionYuxuan Li, Xiang Li, Yunheng Li, Yicheng Zhang 等AAAI 2026 · 被引用 25 次
- XTrack: Multimodal Training Boosts RGB-X Video Object TrackersYuedong Tan, Zongwei Wu, Yuqian Fu, Zhuyun Zhou 等ICCV 2025 · 被引用 10 次
- Sparse Models, Sparse Safety: Unsafe Routes in Mixture-of-Experts LLMsYukun Jiang, Hai Huang, Mingjie Li, Yage Zhang 等ICML 2026 · 被引用 9 次
- Point-MoE: Large-Scale Multi-Dataset Training with Mixture-of-Experts for 3D Semantic SegmentationXuweiyi Chen, Wentao Zhou, Aruni RoyChowdhury, Zezhou ChengICLR 2026 · 被引用 4 次
- COME: Dual Structure-Semantic Learning with Collaborative MOE for Universal Lesion Detection Across Heterogeneous Ultrasound DatasetsLingyu Chen, Yawen Zeng, Yue Wang, Peng Wan 等ICCV 2025 · 被引用 1 次
它引用的顶会 Paper15
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen 等ICLR 2021 · 被引用 1,954 次
- DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETRShilong Liu, Feng Li, Hao Zhang, Xiao Yang 等ICLR 2022 · 被引用 1,218 次
- Scaling Vision with Sparse Mixture of ExpertsCarlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann 等NeurIPS 2021 · 被引用 1,213 次
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong 等ICML 2022 · 被引用 1,173 次
- DN-DETR: Accelerate DETR Training by Introducing Query DeNoisingFeng Li, Hao Zhang, Shilong Liu, Jian Guo 等CVPR 2022 · 被引用 879 次
相关 Paper
- Merging Multi-Task Models via Weight-Ensembling Mixture of ExpertsAnke Tang, Li Shen, Yong Luo, Nan Yin 等ICML 2024 · 被引用 96 次
- Object-Aware Domain Generalization for Object DetectionWooju Lee, Dasol Hong, Hyungtae Lim, Hyun MyungAAAI 2024 · 被引用 58 次
- DON'T NEED RETRAINING: A Mixture of DETR and Vision Foundation Models for Cross-Domain Few-Shot Object DetectionChanghan Liu, Xunzhi Xiang, Zixuan Duan, Wenbin Li 等NeurIPS 2025 · 被引用 8 次
- AdaMV-MoE: Adaptive Multi-Task Vision Mixture-of-ExpertsTianlong Chen, Xuxi Chen, Xianzhi Du, Abdullah Rashwan 等ICCV 2023 · 被引用 119 次
- AnomalyMoE: Towards a Language-free Generalist Model for Unified Visual Anomaly DetectionZhaopeng Gu, Bingke Zhu, Guibo Zhu, Yingying Chen 等AAAI 2026 · 被引用 3 次
