CrackFormer: Transformer Network for Fine-Grained Crack Detection
Huajun Liu, Xiangyu Miao, Christoph Mertz, Chengzhong Xu, Hui Kong
Abstract
Cracks are irregular line structures that are of interest in many computer vision applications. Crack detection (e.g., from pavement images) is a challenging task due to intensity in-homogeneity, topology complexity, low contrast and noisy background. The overall crack detection accuracy can be significantly affected by the detection performance on fine-grained cracks. In this work, we propose a Crack Transformer network (CrackFormer) for fine-grained crack detection. The CrackFormer is composed of novel attention modules in a SegNet-like encoder-decoder architecture. Specifically, it consists of novel self-attention modules with 1x1 convolutional kernels for efficient contextual information extraction across feature-channels, and efficient positional embedding to capture large receptive field contextual information for long range interactions. It also introduces new scaling-attention modules to combine outputs from the corresponding encoder and decoder blocks to suppress nonsemantic features and sharpen semantic ones. The Crack-Former is trained and evaluated on three classical crack datasets. The experimental results show that the Crack-Former achieves the Optimal Dataset Scale (ODS) values of 0.871, 0.877 and 0.881, respectively, on the three datasets and outperforms the state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3a514132-87cc-40bb-b02f-e4d28fb7e2e4Cited by top-tier papers10
- The Devil is in the Crack Orientation: A New Perspective for Crack DetectionZhuangzhuang Chen, Jin Zhang, Zhuonan Lai, Guanming Zhu et al.ICCV 2023 · 26 citations
- MixerCSeg: An Efficient Mixer Architecture for Crack Segmentation via Decoupled Mamba AttentionZilong Zhao, Zhengming Ding, Pei Niu, Wenhao Sun et al.CVPR 2026 · 12 citations
- TopoTTA: Topology-Enhanced Test-Time Adaptation for Tubular Structure SegmentationJiale Zhou, Wenhan Wang, Shikun Li, Xiaolei Qu et al.ICCV 2025 · 2 citations
- LIDAR: Lightweight Adaptive Cue-Aware Fusion Vision Mamba for Multimodal Segmentation of Structural CracksHui Liu, Chen Jia, Fan Shi, Xu Cheng et al.ACM MM 2025 · 1 citation
- Wavelet and Prototype Augmented Query-based Transformer for Pixel-level Surface Defect DetectionFeng Yan, Xiaoheng Jiang, Yang Lu, Jiale Cao et al.CVPR 2025
Builds on4
- CCNet: Criss-Cross Attention for Semantic SegmentationZilong Huang, Xinggang Wang, Lichao Huang, Chang Huang et al.ICCV 2019 · 2,972 citations
- Attention Augmented Convolutional NetworksIrwan Bello, Barret Zoph, Quoc Le, Ashish Vaswani et al.ICCV 2019 · 1,149 citations
- On the Relationship between Self-Attention and Convolutional LayersJean-Baptiste Cordonnier, Andreas Loukas, Martin JaggiICLR 2020 · 629 citations
- LambdaNetworks: Modeling long-range Interactions without AttentionIrwan BelloICLR 2021 · 48 citations
Related papers
- Laneformer: Object-Aware Row-Column Transformers for Lane DetectionJianhua Han, Xiajun Deng, Xinyue Cai, Zhen Yang et al.AAAI 2022 · 79 citations
- Rethinking Semantic Segmentation From a Sequence-to-Sequence Perspective With TransformersSixiao Zheng, Jiachen Lu, Hengshuang Zhao, Xiatian Zhu et al.CVPR 2021
- RAMS-Trans: Recurrent Attention Multi-scale Transformer for Fine-grained Image RecognitionYunqing Hu, Xuan Jin, Yin Zhang, Haiwen Hong et al.ACM MM 2021 · 142 citations
- PEM: Prototype-Based Efficient MaskFormer for Image SegmentationNiccolò Cavagnero, Gabriele Rosi, Claudia Cuttano, Francesca Pistilli et al.CVPR 2024
- T-former: An Efficient Transformer for Image InpaintingYe Deng, Siqi Hui, Sanping Zhou, Deyu Meng et al.ACM MM 2022 · 58 citations
