Towards Open-Vocabulary Industrial Defect Understanding with a Large-Scale Multimodal Dataset
Tsai-Ching Ni, ZhenQi Chen, YuanFu Yang
摘要
We present IMDD-1M, the first large-scale Industrial Multimodal Defect Dataset comprising 1,000,000 aligned image-text pairs, designed to advance multimodal learning for manufacturing and quality inspection. IMDD-1M contains high-resolution real-world defects spanning over 60 material categories and more than 400 defect types, each accompanied by expert-verified annotations and fine-grained textual descriptions detailing defect location, severity, and contextual attributes. This dataset enables a wide spectrum of applications, including classification, segmentation, retrieval, captioning, and generative modeling. Building upon IMDD-1M, we train a diffusion-based vision-language foundation model from scratch, specifically tailored for industrial scenarios. The model serves as a generalizable foundation that can be efficiently adapted to specialized domains through lightweight fine-tuning. With less than 5% of the task-specific data required by dedicated expert models, it achieves comparable performance, highlighting the potential of data-efficient foundation model adaptation for industrial inspection and generation, paving the way for scalable, domain-adaptive, and knowledge-grounded manufacturing intelligence.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- MuSc: Zero-Shot Industrial Anomaly Classification and Segmentation with Mutual Scoring of the Unlabeled ImagesXurui Li, Ziming Huang, Feng Xue, Yu ZhouICLR 2024 · 被引用 76 次
- Masked-attention Mask Transformer for Universal Image SegmentationBowen Cheng, Ishan Misra, Alexander G. Schwing, Alexander Kirillov 等CVPR 2022
相关 Paper
- MMAD: A Comprehensive Benchmark for Multimodal Large Language Models in Industrial Anomaly DetectionXi Jiang, Jian Li, Hanqiu Deng, Yong Liu 等ICLR 2025 · 被引用 3 次
- Zero-Shot Learning in Industrial Scenarios: New Large-Scale Benchmark, Challenges and BaselineZekai Zhang, Qinghui Chen, Maomao Xiong, Shijiao Ding 等AAAI 2025 · 被引用 4 次
- Automated Defect Report Generation for Enhanced Industrial Quality ControlJiayuan Xie, Zhiping Zhou, Zihan Wu, Xinting Zhang 等AAAI 2024 · 被引用 2 次
- FoundIR: Unleashing Million-Scale Training Data to Advance Foundation Models for Image RestorationHao Li, Xiang Chen, Jiangxin Dong, Jinhui Tang 等ICCV 2025 · 被引用 15 次
- Omni-AD: A Large-scale and Versatile Benchmark for Industrial Anomaly DetectionDahu Shi, Chengshen He, Shaochen Zhang, Bo Qian 等CVPR 2026
