DD-LIVM: Pioneering Cross-Domain Photovoltaic Defect Detection Using Large Infrared-Visible Model
Yinan Zhu, Meng Xue, Haiyan Hu, Cong Zhang, Xiaoyi Fan, Qian Zhang
摘要
Photovoltaic (PV) defect detection is crucial for preventing power efficiency loss and fire hazards. The industry primarily relies on the fusion of infrared and visible images for defect localization and diagnosis. However, current detection methods exhibit poor generalizability in new site environments or with altered imaging setups. While recent infrared and vision foundation models (FM) facilitate domain-invariant feature maps extraction, directly concatenating them and fine-tuning achieves limited generalizability gain to PV defect detection, due to the asymmetric dual-modal semantics of defects. In this paper, we present the first large infrared-visible model DD-LIVM to enable cross-domain defect detection. The key innovation of DD-LIVM lies in its defect-specific three-step fine-tuning strategy, which utilizes alternating modality masking. Prior to feature fusion and joint fine-tuning, the infrared and visible FM encoders are alternately masked and optimized to enhance their individual semantic utility for defect localization visibility and classification granularity, with feature distances among different defect types regulated through contrastive learning. This approach allows for the extraction of generalizable and defect-specific feature maps. Moreover, for practical employment of DD-LIVM, we propose a domain-agnostic spatial alignment algorithm for infrared-visible images before dual-modal fusion, and develop source data augmentation and adaptive detection head selection schemes based on defects' infrared characteristics to further enhance the generalizability. Extensive experiments on 7,078 dual-modal images from 9 real-world scenarios across 4 cities' PV stations demonstrate that DD-LIVM achieves an accuracy of 87.7% for cross-domain defect detection, surpassing state-of-the-art methods by 17.3%.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Empowering Visible-Infrared Person Re-Identification with Large Foundation ModelsZhangyi Hu, Bin Yang, Mang YeNeurIPS 2024 · 被引用 45 次
- Dispel Darkness for Better Fusion: A Controllable Visual Enhancer Based on Cross-Modal Conditional Adversarial LearningHao Zhang, Linfeng Tang, Xinyu Xiang, Xuhui Zuo 等CVPR 2024 · 被引用 21 次
- DuGI-MAE: Improving Infrared Mask Autoencoders via Dual-Domain GuidanceYinghui Xing, Xiaoting Su, Shizhou Zhang, Donghao Chu 等AAAI 2026
- DLDA: Unified Dual-Level Domain Adaptation for Low-Light Object DetectionJiayi Hu, Qian Zhao, Gang LiAAAI 2026
- Domain Adaptation Guided Infrared and Visible Image FusionTianwei Guan, Haozhen Wei, Yuhan Zhou, Jun Ma 等AAAI 2026
