DIANet: Dense-and-Implicit Attention Network
Zhongzhan Huang, Senwei Liang, Mingfu Liang, Haizhao Yang
Abstract
Attention networks have successfully boosted the performance in various vision problems. Previous works lay emphasis on designing a new attention module and individually plug them into the networks. Our paper proposes a novel-and-simple framework that shares an attention module throughout different network layers to encourage the integration of layer-wise information and this parameter-sharing module is referred to as Dense-and-Implicit-Attention (DIA) unit. Many choices of modules can be used in the DIA unit. Since Long Short Term Memory (LSTM) has a capacity of capturing long-distance dependency, we focus on the case when the DIA unit is the modified LSTM (called DIA-LSTM). Experiments on benchmark datasets show that the DIA-LSTM unit is capable of emphasizing layer-wise feature interrelation and leads to significant improvement of image classification accuracy. We further empirically show that the DIA-LSTM has a strong regularization ability on stabilizing the training of deep networks by the experiments with the removal of skip connections (He et al. 2016a) or Batch Normalization (Ioffe and Szegedy 2015) in the whole residual network.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9726b76f-e874-4ef5-b7ef-85111515592cCited by top-tier papers12
- Rethinking the Pruning Criteria for Convolutional Neural NetworkZhongzhan Huang, Wenqi Shao, Xinjiang Wang, Liang Lin et al.NeurIPS 2021 · 75 citations
- SUR-adapter: Enhancing Text-to-Image Pre-trained Diffusion Models with Large Language ModelsShanshan Zhong, Zhongzhan Huang, Wushao Wen, Jinghui Qin et al.ACM MM 2023 · 45 citations
- ScaleLong: Towards More Stable Training of Diffusion Model via Scaling Network Long Skip ConnectionZhongzhan Huang, Pan Zhou, Shuicheng Yan, Liang LinNeurIPS 2023 · 41 citations
- Understanding Self-attention Mechanism via Dynamical System PerspectiveZhongzhan Huang, Mingfu Liang, Jinghui Qin, Shanshan Zhong et al.ICCV 2023 · 38 citations
- Mirror Gradient: Towards Robust Multimodal Recommender Systems via Exploring Flat Local MinimaShanshan Zhong, Zhongzhan Huang, Daifeng Li, Wushao Wen et al.WWW 2024 · 24 citations
Related papers
- Cross-Layer Retrospective Retrieving via Layer AttentionYanwen Fang, Yuxi Cai, Jintai Chen, Jingyu Zhao et al.ICLR 2023 · 1 citation
- Recurrence along Depth: Deep Convolutional Neural Networks with Recurrent Layer AggregationJingyu Zhao, Yanwen Fang, Guodong LiNeurIPS 2021 · 31 citations
- Loss-Based Attention for Deep Multiple Instance LearningXiaoshuang Shi, Fuyong Xing, Yuanpu Xie, Zizhao Zhang et al.AAAI 2020 · 123 citations
- AA-RMVSNet: Adaptive Aggregation Recurrent Multi-view Stereo NetworkZizhuang Wei, Qingtian Zhu, Chen Min, Yisong Chen et al.ICCV 2021 · 193 citations
- AttentionRNN: A Structured Spatial Attention MechanismSiddhesh Khandelwal, Leonid SigalICCV 2019 · 4 citations
