Wavelet and Prototype Augmented Query-based Transformer for Pixel-level Surface Defect Detection
Feng Yan, Xiaoheng Jiang, Yang Lu, Jiale Cao, Dong Chen, Mingliang Xu
Abstract
As an important part of intelligent manufacturing, pixellevel surface defect detection (SDD) aims to locate defect areas through mask prediction. Previous methods adopt the image-independent static convolution to indiscriminately classify per-pixel features for mask prediction, which leads to suboptimal results for some challenging scenes such as weak defects and cluttered backgrounds. In this paper, inspired by query-based methods, we propose a Wavelet and Prototype Augmented Query-based Transformer (WP-Former) for surface defect detection. Specifically, a set of dynamic queries for mask prediction is updated through the dual-domain transformer decoder. Firstly, a Waveletenhanced Cross-Attention (WCA) is proposed, which aggregates meaningful high-and low-frequency information of image features in the wavelet domain to refine queries. WCA enhances the representation of high-frequency components by capturing multi-scale relationships between different frequency components, enabling queries to focus more on defect details. Secondly, a Prototype-guided Cross-Attention (PCA) is proposed to refine queries through metaprototypes in the spatial domain. The prototypes aggregate semantically meaningful tokens from image features, facilitating queries to aggregate crucial defect information under the cluttered backgrounds. Extensive experiments on three defect detection datasets (i.e., ESDIs-SOD, CrackSeg9k, and ZJU-Leaper) demonstrate that the proposed method achieves state-of-the-art performance in defect detection. The code will be available at https: //github.com/yfhdm/WPFormer .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 347d7770-5137-4dba-bd88-0e96e8dc4a28Cited by top-tier papers1
Ask how each one uses itBuilds on17
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 2,196 citations
- FcaNet: Frequency Channel Attention NetworksZequn Qin, Pengyi Zhang, Fei Wu, Xi LiICCV 2021 · 1,049 citations
- EMCAD: Efficient Multi-Scale Convolutional Attention Decoding for Medical Image SegmentationMd Mostafijur Rahman, Mustafa Munir, Radu MarculescuCVPR 2024 · 352 citations
- Detecting Camouflaged Object in Frequency DomainYijie Zhong, Bo Li, Lv Tang, Senyun Kuang et al.CVPR 2022 · 271 citations
- CrackFormer: Transformer Network for Fine-Grained Crack DetectionHuajun Liu, Xiangyu Miao, Christoph Mertz, Chengzhong Xu et al.ICCV 2021 · 195 citations
Related papers
- Dynamic Focus-aware Positional Queries for Semantic SegmentationHaoyu He, Jianfei Cai, Zizheng Pan, Jing Liu et al.CVPR 2023
- Pixel-Level Anomaly Detection via Uncertainty-aware Prototypical TransformerChao Huang, Chengliang Liu, Zheng Zhang, Zhihao Wu et al.ACM MM 2022 · 28 citations
- MGQFormer: Mask-Guided Query-Based Transformer for Image Manipulation LocalizationKunlun Zeng, Ri Cheng, Weimin Tan, Bo YanAAAI 2024 · 23 citations
- Scaling up Image Segmentation across Data and TasksPei Wang, Zhaowei Cai, Hao Yang, Ashwin Swaminathan et al.CVPR 2025
- Point Cloud Semantic Scene Completion with Prototype-Guided TransformerChenghao Fang, Jianqing Liang, Jiye Liang, Zijin Du et al.AAAI 2026
