Wavelet and Prototype Augmented Query-based Transformer for Pixel-level Surface Defect Detection
Feng Yan, Xiaoheng Jiang, Yang Lu, Jiale Cao, Dong Chen, Mingliang Xu
摘要
As an important part of intelligent manufacturing, pixellevel surface defect detection (SDD) aims to locate defect areas through mask prediction. Previous methods adopt the image-independent static convolution to indiscriminately classify per-pixel features for mask prediction, which leads to suboptimal results for some challenging scenes such as weak defects and cluttered backgrounds. In this paper, inspired by query-based methods, we propose a Wavelet and Prototype Augmented Query-based Transformer (WP-Former) for surface defect detection. Specifically, a set of dynamic queries for mask prediction is updated through the dual-domain transformer decoder. Firstly, a Waveletenhanced Cross-Attention (WCA) is proposed, which aggregates meaningful high-and low-frequency information of image features in the wavelet domain to refine queries. WCA enhances the representation of high-frequency components by capturing multi-scale relationships between different frequency components, enabling queries to focus more on defect details. Secondly, a Prototype-guided Cross-Attention (PCA) is proposed to refine queries through metaprototypes in the spatial domain. The prototypes aggregate semantically meaningful tokens from image features, facilitating queries to aggregate crucial defect information under the cluttered backgrounds. Extensive experiments on three defect detection datasets (i.e., ESDIs-SOD, CrackSeg9k, and ZJU-Leaper) demonstrate that the proposed method achieves state-of-the-art performance in defect detection. The code will be available at https: //github.com/yfhdm/WPFormer .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper17
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 被引用 2,196 次
- FcaNet: Frequency Channel Attention NetworksZequn Qin, Pengyi Zhang, Fei Wu, Xi LiICCV 2021 · 被引用 1,049 次
- EMCAD: Efficient Multi-Scale Convolutional Attention Decoding for Medical Image SegmentationMd Mostafijur Rahman, Mustafa Munir, Radu MarculescuCVPR 2024 · 被引用 352 次
- Detecting Camouflaged Object in Frequency DomainYijie Zhong, Bo Li, Lv Tang, Senyun Kuang 等CVPR 2022 · 被引用 271 次
- CrackFormer: Transformer Network for Fine-Grained Crack DetectionHuajun Liu, Xiangyu Miao, Christoph Mertz, Chengzhong Xu 等ICCV 2021 · 被引用 195 次
相关 Paper
- Dynamic Focus-aware Positional Queries for Semantic SegmentationHaoyu He, Jianfei Cai, Zizheng Pan, Jing Liu 等CVPR 2023
- Pixel-Level Anomaly Detection via Uncertainty-aware Prototypical TransformerChao Huang, Chengliang Liu, Zheng Zhang, Zhihao Wu 等ACM MM 2022 · 被引用 28 次
- MGQFormer: Mask-Guided Query-Based Transformer for Image Manipulation LocalizationKunlun Zeng, Ri Cheng, Weimin Tan, Bo YanAAAI 2024 · 被引用 23 次
- Scaling up Image Segmentation across Data and TasksPei Wang, Zhaowei Cai, Hao Yang, Ashwin Swaminathan 等CVPR 2025
- Point Cloud Semantic Scene Completion with Prototype-Guided TransformerChenghao Fang, Jianqing Liang, Jiye Liang, Zijin Du 等AAAI 2026
