Learning Hierarchical Image Segmentation For Recognition and By Recognition
Tsung-Wei Ke, Sangwoo Mo, Stella X. Yu
摘要
Large vision and language models learned directly through image-text associations often lack detailed visual substantiation, whereas image segmentation tasks are treated separately from recognition, supervisedly learned without interconnections. Our key observation is that, while an image can be recognized in multiple ways, each has a consistent part-and-whole visual organization. Segmentation thus should be treated not as an end task to be mastered through supervised learning, but as an internal process that evolves with and supports the ultimate goal of recognition. We propose to integrate a hierarchical segmenter into the recognition process, train and adapt the entire model solely on image-level recognition objectives. We learn hierarchical segmentation for free alongside recognition, automatically uncovering part-to-whole relationships that not only underpin but also enhance recognition. Enhancing the Vision Transformer (ViT) with adaptive segment tokens and graph pooling, our model surpasses ViT in unsupervised part-whole discovery, semantic segmentation, image classification, and efficiency. Notably, our model (trained on unlabeled 1M ImageNet images) outperforms SAM (trained on 11M images and 1 billion masks) by absolute 8% in mIoU on PartImageNet object segmentation. * Equal contribution. Code available at https://github.com/twke18/CAST .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- RACCooN: Versatile Instructional Video Editing with Auto-Generated NarrativesJaehong Yoon, Shoubin Yu, Mohit BansalEMNLP 2025 · 被引用 2 次
- Native Segmentation Vision TransformersGuillem Brasó, Aljosa Osep, Laura Leal-TaixéNeurIPS 2025 · 被引用 2 次
- Knowledge-Guided Part SegmentationXuejian Gou, Fang Liu, Licheng Jiao, Shuo Li 等ICCV 2025 · 被引用 1 次
- Open Ad-hoc Categorization with Contextualized Feature LearningZilin Wang, Sangwoo Mo, Stella X. Yu, Sima Behpour 等CVPR 2025
- Improving Visual Recognition with Hyperbolical Visual Hierarchy MappingHyeongjun Kwon, Jinhyun Jang, Jin Kim, Kwonyoung Kim 等CVPR 2024
它引用的顶会 Paper39
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
相关 Paper
- GroupViT: Semantic Segmentation Emerges from Text SupervisionJiarui Xu, Shalini De Mello, Sifei Liu, Wonmin Byeon 等CVPR 2022 · 被引用 398 次
- Open-Vocabulary Universal Image Segmentation with MaskCLIPZheng Ding, Jieke Wang, Zhuowen TuICML 2023 · 被引用 150 次
- The Missing Point in Vision Transformers for Universal Image SegmentationSajjad Shahabodini, Mobina Mansoori, Farnoush Bayatmakou, Jamshid Abouei 等CVPR 2026 · 被引用 5 次
- Segmenter: Transformer for Semantic SegmentationRobin Strudel, Ricardo Garcia, Ivan Laptev, Cordelia SchmidICCV 2021 · 被引用 1,898 次
- Self-Supervised Learning of Object Parts for Semantic SegmentationAdrian Ziegler, Yuki M. AsanoCVPR 2022 · 被引用 87 次
