Top-Down Visual Attention from Analysis by Synthesis
Baifeng Shi, Trevor Darrell, Xin Wang
Abstract
Current attention algorithms (e.g., self-attention) are stimulus-driven and highlight all the salient objects in an image. However, intelligent agents like humans often guide their attention based on the high-level task at hand, focusing only on task-related objects. This ability of taskguided top-down attention provides task-adaptive representation and helps the model generalize to various tasks. In this paper, we consider top-down attention from a classic Analysis-by-Synthesis (AbS) perspective of vision. Prior work indicates a functional equivalence between visual attention and sparse reconstruction; we show that an AbS visual system that optimizes a similar sparse reconstruction objective modulated by a goal-directed top-down signal naturally simulates top-down attention. We further propose Analysis-by-Synthesis Vision Transformer (AbSViT), which is a top-down modulated ViT model that variationally approximates AbS, and achieves controllable top-down attention. For real-world applications, AbSViT consistently improves over baselines on Vision-Language tasks such as VQA and zero-shot retrieval where language guides the top-down attention. AbSViT can also serve as a general backbone, improving performance on classification, semantic segmentation, and model robustness. Project page: https://sites.google.com/view/absvit .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers13
- Frozen Transformers in Language Models Are Effective Visual Encoder LayersZiqi Pang, Ziyang Xie, Yunze Man, Yu-Xiong WangICLR 2024 · 54 citations
- Learning Hierarchical Image Segmentation For Recognition and By RecognitionTsung-Wei Ke, Sangwoo Mo, Stella X. YuICLR 2024 · 20 citations
- Learning from Observer Gaze: Zero-Shot Attention Prediction Oriented by Human-Object Interaction RecognitionYuchen Zhou, Linkai Liu, Chao GouCVPR 2024 · 13 citations
- Bootstrapping Top-down Information for Self-modulating Slot AttentionDongwon Kim, Seoyeon Kim, Suha KwakNeurIPS 2024 · 7 citations
- Unlocking the Capabilities of Masked Generative Models for Image Synthesis via Self-GuidanceJiwan Hur, Dong-Jae Lee, Gyojin Han, Jaehyun Choi et al.NeurIPS 2024 · 6 citations
Builds on22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
Related papers
- BUS : Efficient and Effective Vision-language Pre-training with Bottom-Up Patch SummarizationChaoya Jiang, Haiyang Xu, Wei Ye, Qinghao Ye et al.ICCV 2023 · 9 citations
- Question Aware Vision Transformer for Multimodal ReasoningRoy Ganz, Yair Kittenplon, Aviad Aberdam, Elad Ben-Avraham et al.CVPR 2024
- Attention Guided CAM: Visual Explanations of Vision Transformer Guided by Self-AttentionSaebom Leem, Hyunseok SeoAAAI 2024 · 40 citations
- Unifying Top-Down and Bottom-Up Scanpath Prediction Using TransformersZhibo Yang, Sounak Mondal, Seoyoung Ahn, Ruoyu Xue et al.CVPR 2024
- H-ViT: A Hierarchical Vision Transformer for Deformable Image RegistrationMorteza Ghahremani, Mohammad Khateri, Bailiang Jian, Benedikt Wiestler et al.CVPR 2024
