Towards Open-Vocabulary Industrial Defect Understanding with a Large-Scale Multimodal Dataset
Tsai-Ching Ni, ZhenQi Chen, YuanFu Yang
Abstract
We present IMDD-1M, the first large-scale Industrial Multimodal Defect Dataset comprising 1,000,000 aligned image-text pairs, designed to advance multimodal learning for manufacturing and quality inspection. IMDD-1M contains high-resolution real-world defects spanning over 60 material categories and more than 400 defect types, each accompanied by expert-verified annotations and fine-grained textual descriptions detailing defect location, severity, and contextual attributes. This dataset enables a wide spectrum of applications, including classification, segmentation, retrieval, captioning, and generative modeling. Building upon IMDD-1M, we train a diffusion-based vision-language foundation model from scratch, specifically tailored for industrial scenarios. The model serves as a generalizable foundation that can be efficiently adapted to specialized domains through lightweight fine-tuning. With less than 5% of the task-specific data required by dedicated expert models, it achieves comparable performance, highlighting the potential of data-efficient foundation model adaptation for industrial inspection and generation, paving the way for scalable, domain-adaptive, and knowledge-grounded manufacturing intelligence.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e2974fae-b3e3-4ba6-a060-be69a8c77254Builds on7
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- MuSc: Zero-Shot Industrial Anomaly Classification and Segmentation with Mutual Scoring of the Unlabeled ImagesXurui Li, Ziming Huang, Feng Xue, Yu ZhouICLR 2024 · 76 citations
- Masked-attention Mask Transformer for Universal Image SegmentationBowen Cheng, Ishan Misra, Alexander G. Schwing, Alexander Kirillov et al.CVPR 2022
Related papers
- MMAD: A Comprehensive Benchmark for Multimodal Large Language Models in Industrial Anomaly DetectionXi Jiang, Jian Li, Hanqiu Deng, Yong Liu et al.ICLR 2025 · 3 citations
- Zero-Shot Learning in Industrial Scenarios: New Large-Scale Benchmark, Challenges and BaselineZekai Zhang, Qinghui Chen, Maomao Xiong, Shijiao Ding et al.AAAI 2025 · 4 citations
- Automated Defect Report Generation for Enhanced Industrial Quality ControlJiayuan Xie, Zhiping Zhou, Zihan Wu, Xinting Zhang et al.AAAI 2024 · 2 citations
- FoundIR: Unleashing Million-Scale Training Data to Advance Foundation Models for Image RestorationHao Li, Xiang Chen, Jiangxin Dong, Jinhui Tang et al.ICCV 2025 · 15 citations
- Omni-AD: A Large-scale and Versatile Benchmark for Industrial Anomaly DetectionDahu Shi, Chengshen He, Shaochen Zhang, Bo Qian et al.CVPR 2026
