Zero-Shot Learning in Industrial Scenarios: New Large-Scale Benchmark, Challenges and Baseline
Zekai Zhang, Qinghui Chen, Maomao Xiong, Shijiao Ding, Zhanzhi Su, Xinjie Yao, Yiming Sun, Cong Bai, Jinglin Zhang
Abstract
Large Visual Language Models (LVLMs) have achieved remarkable success in vision tasks. However, the significant differences between industrial and natural scenes make applying LVLMs challenging. Existing LVLMs rely on user-provided prompts to segment objects. This often leads to suboptimal performance due to the inclusion of irrelevant pixels. In addition, the scarcity of data also makes the application of LVLMs in industrial scenarios remain unexplored. To fill this gap, this paper proposes an open industrial dataset and a Refined Text-Visual Prompt (RTVP) for zero-shot industrial defect detection. First, this paper constructs the Multi-Modal Industrial Open Dataset (MMIO) containing 80K+ samples. MMIO contains diverse industrial categories, including 6 super categories and 18 subcategories. MMIO is the first large-scale multi-scenes pre-training dataset for industrial zero-shot learning, and provides valuable training data for open models in future industrial scenarios. Based on MMIO, this paper provides a RTVP specifically for industrial zero-shot tasks. RTVP has two significant advantages: First, this paper designs an expert-guided large model domain adaptation mechanism and designs an industrial zero-shot method based on Mobile-SAM, which enhances the generalization ability of large models in industrial scenarios. Second, RTVP automatically generates visual prompts directly from images and considers text-visual prompt interactions ignored by previous LVLM, improving visual and textual content understanding. RTVP achieves SOTA with 42.2% and 24.7% AP in zero-shot and closed scenes of MMIO.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Back to Point: Exploring Point-Language Models for Zero-Shot 3D Anomaly DetectionKaiqiang Li, Gang Li, Mingle Zhou, Min Li et al.CVPR 2026 · 2 citations
- KFTD: Koopman-Fourier Time-Differentiable Network for Continuous Ocean Spatiotemporal ForecastingQinghui Chen, Zekai Zhang, Hailong Liu, Jinglin Zhang et al.KDD 2026
- ADSeeker: A Knowledge-Grounded Reasoning Framework for Industry Anomaly Detection and ReasoningKai Zhang, Zekai Zhang, Xihe Sun, Anpeng Wang et al.CVPR 2026
Builds on18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- YOLOv10: Real-Time End-to-End Object DetectionAo Wang, Hui Chen, Lihao Liu, Kai Chen et al.NeurIPS 2024 · 6,113 citations
Related papers
- Aligning and Prompting Anything for Zero-Shot Generalized Anomaly DetectionJitao Ma, Weiying Xie, Hangyu Ye, Daixun Li et al.AAAI 2025 · 3 citations
- MMAD: A Comprehensive Benchmark for Multimodal Large Language Models in Industrial Anomaly DetectionXi Jiang, Jian Li, Hanqiu Deng, Yong Liu et al.ICLR 2025 · 3 citations
- Fine-Grained Abnormality Prompt Learning for Zero-Shot Anomaly DetectionJiawen Zhu, Yew-Soon Ong, Chunhua Shen, Guansong PangICCV 2025 · 14 citations
- AnomalyGPT: Detecting Industrial Anomalies Using Large Vision-Language ModelsZhaopeng Gu, Bingke Zhu, Guibo Zhu, Yingying Chen et al.AAAI 2024 · 312 citations
- PromptMoE: Generalizable Zero-Shot Anomaly Detection via Visually-Guided Prompt MixturesYuheng Shao, Lizhang Wang, Changhao Li, Peixian Chen et al.AAAI 2026
