Finding Distributed Object-Centric Properties in Self-Supervised Transformers
Samyak Rawlekar, Amitabh Swain, Yujun Cai, Yiwei Wang, Ming-Hsuan Yang, Narendra Ahuja
摘要
Self-supervised Vision Transformers (ViTs) like DINO show an emergent ability to discover objects, typically observed in [CLS] token attention maps of the final layer. However, these maps often contain spurious activations resulting in poor localization of objects. This is because the [CLS] token, trained on an image-level objective, summarizes the entire image instead of focusing on objects. This aggregation dilutes the object-centric information existing in the local, patch-level interactions. We analyze this by computing inter-patch similarity using patch-level attention components (query, key, and value) across all layers. We find that: (1) Object-centric properties are encoded in the similarity maps derived from all three components (), unlike prior work that uses only key features or the [CLS] token. (2) This object-centric information is distributed across the network, not just confined to the final layer. Based on these insights, we introduce Object-DINO, a training-free method that extracts this distributed object-centric information. Object-DINO clusters attention heads across all layers based on the similarities of their patches and automatically identifies the object-centric cluster corresponding to all objects. We demonstrate Object-DINO's effectiveness on two applications: enhancing unsupervised object discovery (+3.6 to +12.4 CorLoc gains) and mitigating object hallucination in Multimodal Large Language Models by providing visual grounding. Our results demonstrate that using this distributed object-centric information improves downstream tasks without additional training.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
- Evaluating Object Hallucination in Large Vision-Language ModelsYifan Li, Yifan Du, Kun Zhou, Jinpeng Wang 等EMNLP 2023 · 被引用 344 次
- Understanding The Robustness in Vision TransformersDaquan Zhou, Zhiding Yu, Enze Xie, Chaowei Xiao 等ICML 2022 · 被引用 242 次
- HALC: Object Hallucination Reduction via Adaptive Focal-Contrast DecodingZhaorun Chen, Zhuokai Zhao, Hongyin Luo, Huaxiu Yao 等ICML 2024 · 被引用 164 次
相关 Paper
- Patch-level Representation Learning for Self-supervised Vision TransformersSukmin Yun, Hankook Lee, Jaehyung Kim, Jinwoo ShinCVPR 2022 · 被引用 52 次
- MOST: Multiple Object localization with Self-supervised Transformers for object discoverySai Saketh Rambhatla, Ishan Misra, Rama Chellappa, Abhinav ShrivastavaICCV 2023 · 被引用 21 次
- Beyond the Global Scores: Fine-Grained Token Grounding as a Robust Detector of LVLM HallucinationsTuan Dung Nguyen, Minh Khoi Ho, Qi Chen, Yutong Xie 等CVPR 2026 · 被引用 4 次
- Self-Supervised Transformers for Unsupervised Object Discovery using Normalized CutYangtao Wang, Xi Shen, Shell Xu Hu, Yuan Yuan 等CVPR 2022 · 被引用 143 次
- Vision Transformers Need More Than RegistersCheng Shi, Yizhou Yu, Sibei YangCVPR 2026 · 被引用 17 次
