Complementary Prototype Mapping for Efficient Multimodal Anomaly Detection
Yuan Zhao, Zhang xiaoqin to Xiaoqin Zhang, Huchuan Lu, Lihe Zhang
摘要
Multimodal unsupervised anomaly detection has garnered increasing attention for robust defect localization.Recent approaches rely on establishing cross-modal matching relationships under normal conditions without explicit guidance.However, in practice, a single modality may have multiple distinct representations corresponding to another modality, and such unconditional mappings struggle to adaptively capture these variations, resulting in mapping ambiguity and the misclassification of diverse yet normal variations as anomalies.Moreover, existing methods suffer from slow inference speed and high memory overhead, hindering their deployment in real-world production lines.To address these issues, we propose an efficient and effective Complementary Prototype Mapping (CPMAD) framework, which dynamically extracts consensus and supplementary prototypes to serve as complementary priors, thereby guiding and disambiguating cross-modal mappings.The framework comprises three key components:(1) Consensus Extraction Module (CEM) learns a dynamic anchor, transforming multimodal features into anomaly-free consensus prototypes to improve cross-modal consistency and suppress latent anomalies;(2) Supplementary Query Module (SQM) employs a Complementary Residual Attention mechanism to capture the discrepancy between the consensus and modality-specific spaces, thereby exploring the most representative and discriminative cues as supplementary prototypes; and(3) Complementary Mapping Module adaptively integrates both prototypes to perform feature mapping.Extensive experiments demonstrate that CPMAD not only achieves superior performance in both full-data and few-shot settings across diverse industrial and medical scenarios but also maintains faster inference speeds and lower memory consumption compared to existing methods.The code will be released upon publication.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper18
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Towards Total Recall in Industrial Anomaly DetectionKarsten Roth, Latha Pemula, Joaquin Zepeda, Bernhard Schölkopf 等CVPR 2022 · 被引用 1,301 次
- Anomaly Detection via Reverse Distillation from One-Class EmbeddingHanqiu Deng, Xingyu LiCVPR 2022 · 被引用 701 次
- A Unified Model for Multi-class Anomaly DetectionZhiyuan You, Lei Cui, Yujun Shen, Kai Yang 等NeurIPS 2022 · 被引用 585 次
相关 Paper
- PIRN: Prototypical-based Intra-modal Reconstruction with Normality Communication for Multi-modal Anomaly Detection.YITING LI, Xulei Yang, Jing Zhang, Sichao Tian 等ICLR 2026
- Beyond Single-Modal Boundary: Cross-Modal Anomaly Detection through Visual Prototype and HarmonizationKai Mao, Ping Wei, Yiyang Lian, Yangyang Wang 等CVPR 2025
- FAMRD: Frequency-Aware Multimodal Reverse Distillation for Industrial Anomaly DetectionQiyin Zhong, Xianglin Qiu, Xiaolei Wang, Zhen Zhang 等ACM MM 2025
- Multimodal Industrial Anomaly Detection by Crossmodal Feature MappingAlex Costanzino, Pierluigi Zama Ramirez, Giuseppe Lisanti, Luigi Di StefanoCVPR 2024
- Exploring Multimodal Prompts For Unsupervised Continuous Anomaly DetectionMingle Zhou, Jiahui Liu, Jin Wan, Gang Li 等ACM MM 2025
