PIRN: Prototypical-based Intra-modal Reconstruction with Normality Communication for Multi-modal Anomaly Detection.
YITING LI, Xulei Yang, Jing Zhang, Sichao Tian, Jingyi Liao, Fayao Liu
Abstract
Unsupervised multimodal anomaly detection (MAD) aims to detect anomalies by using both RGB and 3D modalities. However, existing methods struggle in few-shot scenarios where the number of normal training samples is limited. Specifically, cross-modal alignment approaches fail to learn reliable correspondences from scarce normal data, whereas memory-based methods often misclassify unseen normal variations as anomalies. To address these issues, we propose , a prototype-driven reconstruction framework equipped with explicit cross-modal knowledge transfer. Instead of relying on dense feature alignment or heavy memory banks, uses a compact set of learnable prototypes to capture diverse normal patterns and constrain feature reconstruction. Specifically, our framework incorporates three core innovations. We introduce Balanced Prototype Assignment (BPA), which employs balanced optimal transport to ensure uniform prototype utilization and prevent codebook collapse. Next, we propose Adaptive Prototype Refinement (APR), which uses gated prototype updates to dynamically expand the model's knowledge of unseen normal variations during inference. To enable each modality to assist the other in reconstructing, we further develop a Multimodal Normality Communication (MNC) module that exchanges high-level normal cues between modalities via gated cross-attention. Extensive experiments on the MVTec 3D-AD, Eyecandies, and Real-IAD benchmarks validate the effectiveness of , where it consistently achieves superior performance compared to existing baselines under challenging few-shot settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on21
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Hierarchical Vector Quantized Transformer for Multi-class Unsupervised Anomaly DetectionRuiying Lu, Yujie Wu, Long Tian, Dongsheng Wang et al.NeurIPS 2023 · 121 citations
- FastRecon: Few-shot Industrial Anomaly Detection via Fast Feature ReconstructionZheng Fang, Xiaoyang Wang, Haocheng Li, Jiejie Liu et al.ICCV 2023 · 100 citations
- Shape-Guided Dual-Memory Learning for 3D Anomaly DetectionYu-Min Chu, Chieh Liu, Ting-I Hsieh, Hwann-Tzong Chen et al.ICML 2023 · 80 citations
- Online Prototype Learning for Online Continual LearningYujie Wei, Jiaxin Ye, Zhizhong Huang, Junping Zhang et al.ICCV 2023 · 78 citations
Related papers
- Remove the Ambiguity: Few-shot Multimodal Anomaly Detection Using Crossmodal Feature ReplacersYuan Guo, Wanqi Zhang, Xu WangICML 2026
- FastRef: Fast Prototype Refinement for Few-shot Industrial Anomaly DetectionYufei Li, Long Tian, Yuyang Dai, Wenchao Chen et al.CVPR 2026 · 7 citations
- Is Task-Specific Training Necessary for Anomaly Detection?Xingwu Zhang, Guanxuan Li, Paul Henderson, Gerardo Aragon-Camarasa et al.ICML 2026 · 1 citation
- Complementary Prototype Mapping for Efficient Multimodal Anomaly DetectionYuan Zhao, Zhang xiaoqin to Xiaoqin Zhang, Huchuan Lu, Lihe ZhangCVPR 2026
- Beyond Single-Modal Boundary: Cross-Modal Anomaly Detection through Visual Prototype and HarmonizationKai Mao, Ping Wei, Yiyang Lian, Yangyang Wang et al.CVPR 2025
