UniMMAD: Unified Multi-Modal and Multi-Class Anomaly Detection via MoE-Driven Feature Decompression
Yuan Zhao, Youwei Pang, Lihe Zhang, Hanqi Liu, Jiaming Zuo, Huchuan Lu, Xiaoqi Zhao
摘要
Existing anomaly detection methods often treat the modality and class as independent factors. Although this paradigm has enriched the development of AD research branches and produced many specialized models, it has also led to fragmented solutions and excessive memory overhead. Moreover, reconstruction-based multi-class approaches typically rely on shared decoding paths, which struggle to handle large variations across domains, resulting in distorted normality boundaries, domain interference, and high false alarm rates. To address these limitations, we propose UniMMAD, a unified framework for multi-modal and multi-class anomaly detection. At the core of UniMMAD is a Mixture-of-Experts (MoE)-driven feature decompression mechanism, which enables adaptive and disentangled reconstruction tailored to specific domains. This process is guided by a "general → specific" paradigm. In the encoding stage, multi-modal inputs of varying combinations are compressed into compact, general-purpose features. The encoder incorporates a feature compression module to suppress latent anomalies, encourage cross-modal interaction, and avoid shortcut learning. In the decoding stage, the general features are decompressed into modality-specific and class-specific forms via a sparsely-gated cross MoE, which dynamically selects expert pathways based on input modality and class. To further improve efficiency, we design a grouped dynamic filtering mechanism and an MoE-in-MoE structure, reducing MoE parameter usage by approximately 75% while maintaining sparse activation and fast inference. UniMMAD achieves state-of-theart performance on 9 anomaly detection datasets, spanning 3 fields, 12 modalities, and 66 classes. Code is publicly available at https://github.com/yuanzhao-CVLAB/UniMMAD.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Towards an Incremental Unified Multimodal Anomaly Detection: Augmenting Multimodal Denoising From an Information Bottleneck PerspectiveKaifang Long, Lianbo Ma, Jiaqi Liu, liming liu 等CVPR 2026 · 被引用 5 次
- Complementary Prototype Mapping for Efficient Multimodal Anomaly DetectionYuan Zhao, Zhang xiaoqin to Xiaoqin Zhang, Huchuan Lu, Lihe ZhangCVPR 2026
它引用的顶会 Paper20
- Towards Total Recall in Industrial Anomaly DetectionKarsten Roth, Latha Pemula, Joaquin Zepeda, Bernhard Schölkopf 等CVPR 2022 · 被引用 1,301 次
- Scaling Vision with Sparse Mixture of ExpertsCarlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann 等NeurIPS 2021 · 被引用 1,213 次
- Anomaly Detection via Reverse Distillation from One-Class EmbeddingHanqiu Deng, Xingyu LiCVPR 2022 · 被引用 701 次
- A Unified Model for Multi-class Anomaly DetectionZhiyuan You, Lei Cui, Yujun Shen, Kai Yang 等NeurIPS 2022 · 被引用 585 次
- PromptIR: Prompting for All-in-One Image RestorationVaishnav Potlapalli, Syed Waqas Zamir, Salman H. Khan, Fahad Shahbaz KhanNeurIPS 2023 · 被引用 386 次
相关 Paper
- AnomalyMoE: Towards a Language-free Generalist Model for Unified Visual Anomaly DetectionZhaopeng Gu, Bingke Zhu, Guibo Zhu, Yingying Chen 等AAAI 2026 · 被引用 3 次
- One-for-All: Proposal Masked Cross-Class Anomaly DetectionXincheng Yao, Chongyang Zhang, Ruoqi Li, Jun Sun 等AAAI 2023 · 被引用 42 次
- Dinomaly: The Less Is More Philosophy in Multi-Class Unsupervised Anomaly DetectionJia Guo, Shuai Lu, Weihang Zhang, Fang Chen 等CVPR 2025
- PIRN: Prototypical-based Intra-modal Reconstruction with Normality Communication for Multi-modal Anomaly Detection.YITING LI, Xulei Yang, Jing Zhang, Sichao Tian 等ICLR 2026
- A Diffusion-Based Framework for Multi-Class Anomaly DetectionHaoyang He, Jiangning Zhang, Hongxu Chen, Xuhai Chen 等AAAI 2024 · 被引用 231 次
