UniMMAD: Unified Multi-Modal and Multi-Class Anomaly Detection via MoE-Driven Feature Decompression
Yuan Zhao, Youwei Pang, Lihe Zhang, Hanqi Liu, Jiaming Zuo, Huchuan Lu, Xiaoqi Zhao
Abstract
Existing anomaly detection methods often treat the modality and class as independent factors. Although this paradigm has enriched the development of AD research branches and produced many specialized models, it has also led to fragmented solutions and excessive memory overhead. Moreover, reconstruction-based multi-class approaches typically rely on shared decoding paths, which struggle to handle large variations across domains, resulting in distorted normality boundaries, domain interference, and high false alarm rates. To address these limitations, we propose UniMMAD, a unified framework for multi-modal and multi-class anomaly detection. At the core of UniMMAD is a Mixture-of-Experts (MoE)-driven feature decompression mechanism, which enables adaptive and disentangled reconstruction tailored to specific domains. This process is guided by a "general → specific" paradigm. In the encoding stage, multi-modal inputs of varying combinations are compressed into compact, general-purpose features. The encoder incorporates a feature compression module to suppress latent anomalies, encourage cross-modal interaction, and avoid shortcut learning. In the decoding stage, the general features are decompressed into modality-specific and class-specific forms via a sparsely-gated cross MoE, which dynamically selects expert pathways based on input modality and class. To further improve efficiency, we design a grouped dynamic filtering mechanism and an MoE-in-MoE structure, reducing MoE parameter usage by approximately 75% while maintaining sparse activation and fast inference. UniMMAD achieves state-of-theart performance on 9 anomaly detection datasets, spanning 3 fields, 12 modalities, and 66 classes. Code is publicly available at https://github.com/yuanzhao-CVLAB/UniMMAD.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Towards an Incremental Unified Multimodal Anomaly Detection: Augmenting Multimodal Denoising From an Information Bottleneck PerspectiveKaifang Long, Lianbo Ma, Jiaqi Liu, liming liu et al.CVPR 2026 · 5 citations
- Complementary Prototype Mapping for Efficient Multimodal Anomaly DetectionYuan Zhao, Zhang xiaoqin to Xiaoqin Zhang, Huchuan Lu, Lihe ZhangCVPR 2026
Builds on20
- Towards Total Recall in Industrial Anomaly DetectionKarsten Roth, Latha Pemula, Joaquin Zepeda, Bernhard Schölkopf et al.CVPR 2022 · 1,301 citations
- Scaling Vision with Sparse Mixture of ExpertsCarlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann et al.NeurIPS 2021 · 1,213 citations
- Anomaly Detection via Reverse Distillation from One-Class EmbeddingHanqiu Deng, Xingyu LiCVPR 2022 · 701 citations
- A Unified Model for Multi-class Anomaly DetectionZhiyuan You, Lei Cui, Yujun Shen, Kai Yang et al.NeurIPS 2022 · 585 citations
- PromptIR: Prompting for All-in-One Image RestorationVaishnav Potlapalli, Syed Waqas Zamir, Salman H. Khan, Fahad Shahbaz KhanNeurIPS 2023 · 386 citations
Related papers
- AnomalyMoE: Towards a Language-free Generalist Model for Unified Visual Anomaly DetectionZhaopeng Gu, Bingke Zhu, Guibo Zhu, Yingying Chen et al.AAAI 2026 · 3 citations
- One-for-All: Proposal Masked Cross-Class Anomaly DetectionXincheng Yao, Chongyang Zhang, Ruoqi Li, Jun Sun et al.AAAI 2023 · 42 citations
- Dinomaly: The Less Is More Philosophy in Multi-Class Unsupervised Anomaly DetectionJia Guo, Shuai Lu, Weihang Zhang, Fang Chen et al.CVPR 2025
- PIRN: Prototypical-based Intra-modal Reconstruction with Normality Communication for Multi-modal Anomaly Detection.YITING LI, Xulei Yang, Jing Zhang, Sichao Tian et al.ICLR 2026
- A Diffusion-Based Framework for Multi-Class Anomaly DetectionHaoyang He, Jiangning Zhang, Hongxu Chen, Xuhai Chen et al.AAAI 2024 · 231 citations
