Look-Into-Object: Self-Supervised Structure Modeling for Object Recognition
Mohan Zhou, Yalong Bai, Wei Zhang, Tiejun Zhao, Tao Mei
摘要
Most object recognition approaches predominantly focus on learning discriminative visual patterns while overlooking the holistic object structure. Though important, structure modeling usually requires significant manual annotations and therefore is labor-intensive. In this paper, we propose to "look into object" (explicitly yet intrinsically model the object structure) through incorporating selfsupervisions into the traditional framework. We show the recognition backbone can be substantially enhanced for more robust representation learning, without any cost of extra annotation and inference speed. Specifically, we first propose an object-extent learning module for localizing the object according to the visual patterns shared among the instances in the same category. We then design a spatial context learning module for modeling the internal structures of the object, through predicting the relative positions within the extent. These two modules can be easily plugged into any backbone networks during training and detached at inference time. Extensive experiments show that our lookinto-object approach (LIO) achieves large performance gain on a number of benchmarks, including generic object recognition (ImageNet) and fine-grained object recognition tasks (CUB, Cars, Aircraft). We also show that this learning paradigm is highly generalizable to other tasks such as object detection and segmentation (MS COCO
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Dual Cross-Attention Learning for Fine-Grained Visual Categorization and Object Re-IdentificationHaowei Zhu, Wenjing Ke, Dong Li, Ji Liu 等CVPR 2022 · 被引用 251 次
- SIM-Trans: Structure Information Modeling Transformer for Fine-grained Visual CategorizationHongbo Sun, Xiangteng He, Yuxin PengACM MM 2022 · 被引用 128 次
- Fine-Grained Object Classification via Self-Supervised Pose AlignmentXuhui Yang, Yaowei Wang, Ke Chen, Yong Xu 等CVPR 2022 · 被引用 91 次
- Joint Representation Learning and Novel Category Discovery on Single- and Multi-modal DataXuhui Jia, Kai Han, Yukun Zhu, Bradley GreenICCV 2021 · 被引用 78 次
- Dynamic Position-aware Network for Fine-grained Image RecognitionShijie Wang, Haojie Li, Zhihui Wang, Wanli OuyangAAAI 2021 · 被引用 36 次
相关 Paper
- Unsupervised Object-Level Representation Learning from Scene ImagesJiahao Xie, Xiaohang Zhan, Ziwei Liu, Yew Soon Ong 等NeurIPS 2021 · 被引用 93 次
- Region Similarity Representation LearningTete Xiao, Colorado J. Reed, Xiaolong Wang, Kurt Keutzer 等ICCV 2021 · 被引用 128 次
- PARTS: Unsupervised segmentation with slots, attention and independence maximizationDaniel Zoran, Rishabh Kabra, Alexander Lerchner, Danilo J. RezendeICCV 2021 · 被引用 53 次
- Unsupervised Part Discovery from Contrastive ReconstructionSubhabrata Choudhury, Iro Laina, Christian Rupprecht, Andrea VedaldiNeurIPS 2021 · 被引用 74 次
- Distilling Localization for Self-Supervised Representation LearningNanxuan Zhao, Zhirong Wu, Rynson W. H. Lau, Stephen LinAAAI 2021 · 被引用 59 次
