PicT: A Slim Weakly Supervised Vision Transformer for Pavement Distress Classification
Wenhao Tang, Sheng Huang, Xiaoxian Zhang, Luwen Huangfu
摘要
Automatic pavement distress classification facilitates improving the efficiency of pavement maintenance and reducing the cost of labor and resources. A recently influential branch of this task divides the pavement image into patches and infers the patch labels for addressing these issues from the perspective of multi-instance learning. However, these methods neglect the correlation between patches and suffer from a low efficiency in the model optimization and inference. As a representative approach of vision Transformer, Swin Transformer is able to address both of these issues. It first provides a succinct and efficient framework for encoding the divided patches as visual tokens, then employs self-attention to model their relations. Built upon Swin Transformer, we present a novel vision Transformer named Pavement Image Classification Transformer (PicT) for pavement distress classification. In order to better exploit the discriminative information of pavement images at the patch level, the Patch Labeling Teacher is proposed to leverage a teacher model to dynamically generate pseudo labels of patches from image labels during each iteration, and guides the model to learn the discriminative features of patches via patch label inference in a weakly supervised manner. The broad classification head of Swin Transformer may dilute the discriminative features of distressed patches in the feature aggregation step due to the small distressed area ratio of the pavement image. To overcome this drawback, we present a Patch Refiner to cluster patches into different groups and only select the highest distress-risk group to yield a slim head for the final image classification. We evaluate our method on a large-scale bituminous pavement distress dataset named CQU-BPDD. Extensive results demonstrate the superiority of our method over baselines and also show that PicT outperforms the second-best performed model by a large margin of +2.4% in [email protected] on detection task, +3.9% in F1 on recognition task, and 1.8x throughput, while enjoying 7x faster training speed using the same computing resources. Our codes and models have been released on https://github.com/DearCaat/PicT.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang 等NeurIPS 2020 · 被引用 5,129 次
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 被引用 3,632 次
相关 Paper
- CrackFormer: Transformer Network for Fine-Grained Crack DetectionHuajun Liu, Xiangyu Miao, Christoph Mertz, Chengzhong Xu 等ICCV 2021 · 被引用 195 次
- Patch-level Representation Learning for Self-supervised Vision TransformersSukmin Yun, Hankook Lee, Jaehyung Kim, Jinwoo ShinCVPR 2022 · 被引用 52 次
- Quantized Feature Distillation for Network QuantizationKe Zhu, Yin-Yin He, Jianxin WuAAAI 2023 · 被引用 21 次
- Not All Images are Worth 16x16 Words: Dynamic Transformers for Efficient Image RecognitionYulin Wang, Rui Huang, Shiji Song, Zeyi Huang 等NeurIPS 2021 · 被引用 283 次
- Faster Vision Transformers with Adaptive PatchesRohan Choudhury, JungEun Kim, Jinhyung Park, Eunho Yang 等ICLR 2026 · 被引用 8 次
