Multimodal Industrial Anomaly Detection by Crossmodal Feature Mapping
Alex Costanzino, Pierluigi Zama Ramirez, Giuseppe Lisanti, Luigi Di Stefano
Abstract
Recent advancements have shown the potential of leveraging both point clouds and images to localize anomalies. Nevertheless, their applicability in industrial manufacturing is often constrained by significant drawbacks, such as the use of memory banks, which leads to a substantial increase in terms of memory footprint and inference times. We propose a novel light and fast framework that learns to map features from one modality to the other on nominal samples and detect anomalies by pinpointing inconsistencies between observed and mapped features. Extensive experiments show that our approach achieves state-of-the-art detection and segmentation performance in both the standard and few-shot settings on the MVTec 3D-AD dataset while achieving faster inference and occupying less memory than previous multimodal AD methods. Furthermore, we propose a layer pruning technique to improve memory and time efficiency with a marginal sacrifice in performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers23
- CNC: Cross-modal Normality Constraint for Unsupervised Multi-class Anomaly DetectionXiaolei Wang, Xiaoyang Wang, Huihui Bai, Eng Gee Lim et al.AAAI 2025 · 19 citations
- Revisiting Multimodal Fusion for 3D Anomaly Detection from an Architectural PerspectiveKaifang Long, Guoyang Xie, Lianbo Ma, Jiaqi Liu et al.AAAI 2025 · 17 citations
- Normal-Abnormal Guided Generalist Anomaly DetectionYuexin Wang, Xiaolei Wang, Yizheng Gong, Jimin XiaoNeurIPS 2025 · 16 citations
- G2SF: Geometry-Guided Score Fusion for Multimodal Industrial Anomaly DetectionChengyu Tao, Xuanming Cao, Juan DuICCV 2025 · 7 citations
- Towards Multimodal Time Series Anomaly Detection with Semantic Alignment and Condensed InteractionShiyan Hu, Jianxin Jin, Yang Shu, Peng Chen et al.ICLR 2026 · 7 citations
Builds on13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Towards Total Recall in Industrial Anomaly DetectionKarsten Roth, Latha Pemula, Joaquin Zepeda, Bernhard Schölkopf et al.CVPR 2022 · 1,301 citations
- Anomaly Detection via Reverse Distillation from One-Class EmbeddingHanqiu Deng, Xingyu LiCVPR 2022 · 701 citations
- Self-Supervised Predictive Convolutional Attentive Block for Anomaly DetectionNicolae-Catalin Ristea, Neelu Madan, Radu Tudor Ionescu, Kamal Nasrollahi et al.CVPR 2022 · 264 citations
Related papers
- Multimodal Industrial Anomaly Detection via Hybrid FusionYue Wang, Jinlong Peng, Jiangning Zhang, Ran Yi et al.CVPR 2023
- FIND: Few-Shot Anomaly Inspection with Normal-Only Multi-Modal DataYiting Li, Fayao Liu, Jingyi Liao, Sichao Tian et al.ICCV 2025 · 5 citations
- Real-IAD D3: A Real-World 2D/Pseudo-3D/3D Dataset for Industrial Anomaly DetectionWenbing Zhu, Lidong Wang, Ziqing Zhou, Chengjie Wang et al.CVPR 2025
- BridgeNet: A Unified Multimodal Framework for Bridging 2D and 3D Industrial Anomaly DetectionAn Xiang, Zixuan Huang, Xitong Gao, Kejiang Ye et al.ACM MM 2025 · 5 citations
- Complementary Prototype Mapping for Efficient Multimodal Anomaly DetectionYuan Zhao, Zhang xiaoqin to Xiaoqin Zhang, Huchuan Lu, Lihe ZhangCVPR 2026
