Learning Dictionary for Visual Attention
Yingjie Liu, Xuan Liu, Hui Yu, Xuan Tang, Xian Wei
Abstract
Recently, the attention mechanism has shown outstanding competence in capturing global structure information and long-range relationships within data, thus enhancing the performance of deep vision models on various computer vision tasks. In this work, we propose a novel alternative dictionary learning-based attention ( Dic-Attn ) module, which models this issue as a decomposition and reconstruction problem with the sparsity prior, inspired by sparse coding in the human visual perception system. The proposed Dic-Attn module decomposes the input into a dictionary and corresponding sparse representations, allowing for the disentanglement of underlying nonlinear structural information in visual data and the reconstruction of an attention embedding. By applying transformation operations in the spatial and channel domains, the module dynamically selects the dictionary’s atoms and sparse representations. Finally, the updated dictionary and sparse representations capture the global contextual information and reconstruct the attention maps. The proposed Dic-Attn module is designed with plug-and-play compatibility and facilitates integration into deep attention encoders. Our approach offers an intuitive and elegant means to exploit the discriminative information from data, promoting visual attention construction. Extensive experimental results on various computer vision tasks, e.g., image and point cloud classification, validate that our method achieves promising performance, and shows a strong competitive comparison with state-of-the-art attention methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e7319445-606c-4850-b3c7-f017d62bf80fBuilds on11
- Revisiting Point Cloud Classification: A New Benchmark Dataset and Classification Model on Real-World DataMikaela Angelina Uy, Quang-Hieu Pham, Binh-Son Hua, Duc Thanh Nguyen et al.ICCV 2019 · 1,003 citations
- Attention on Attention for Image CaptioningLun Huang, Wenmin Wang, Jie Chen, Xiaoyong WeiICCV 2019 · 992 citations
- Point-BERT: Pre-training 3D Point Cloud Transformers with Masked Point ModelingXumin Yu, Lulu Tang, Yongming Rao, Tiejun Huang et al.CVPR 2022 · 684 citations
- Expectation-Maximization Attention Networks for Semantic SegmentationXia Li, Zhisheng Zhong, Jianlong Wu, Yibo Yang et al.ICCV 2019 · 639 citations
- Towards Robust Vision TransformerXiaofeng Mao, Gege Qi, Yuefeng Chen, Xiaodan Li et al.CVPR 2022 · 185 citations
Related papers
- Deep Convolutional Dictionary Learning for Image DenoisingHongyi Zheng, Hongwei Yong, Lei ZhangCVPR 2021
- Learned Image Compression with Dictionary-based Entropy ModelJingbo Lu, Leheng Zhang, Xingyu Zhou, Mu Li et al.CVPR 2025
- In-Context Compositional Learning vis Sparse Coding TransformerWei Chen, Jingxi Yu, Zichen Miao, Qiang QiuNeurIPS 2025
- Pyramid Point Cloud Transformer for Large-Scale Place RecognitionLe Hui, Hang Yang, Mingmei Cheng, Jin Xie et al.ICCV 2021 · 147 citations
- Relation-Aware Global Attention for Person Re-IdentificationZhizheng Zhang, Cuiling Lan, Wenjun Zeng, Xin Jin et al.CVPR 2020
