DIANet: Dense-and-Implicit Attention Network
Zhongzhan Huang, Senwei Liang, Mingfu Liang, Haizhao Yang
摘要
Attention networks have successfully boosted the performance in various vision problems. Previous works lay emphasis on designing a new attention module and individually plug them into the networks. Our paper proposes a novel-and-simple framework that shares an attention module throughout different network layers to encourage the integration of layer-wise information and this parameter-sharing module is referred to as Dense-and-Implicit-Attention (DIA) unit. Many choices of modules can be used in the DIA unit. Since Long Short Term Memory (LSTM) has a capacity of capturing long-distance dependency, we focus on the case when the DIA unit is the modified LSTM (called DIA-LSTM). Experiments on benchmark datasets show that the DIA-LSTM unit is capable of emphasizing layer-wise feature interrelation and leads to significant improvement of image classification accuracy. We further empirically show that the DIA-LSTM has a strong regularization ability on stabilizing the training of deep networks by the experiments with the removal of skip connections (He et al. 2016a) or Batch Normalization (Ioffe and Szegedy 2015) in the whole residual network.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Rethinking the Pruning Criteria for Convolutional Neural NetworkZhongzhan Huang, Wenqi Shao, Xinjiang Wang, Liang Lin 等NeurIPS 2021 · 被引用 75 次
- SUR-adapter: Enhancing Text-to-Image Pre-trained Diffusion Models with Large Language ModelsShanshan Zhong, Zhongzhan Huang, Wushao Wen, Jinghui Qin 等ACM MM 2023 · 被引用 45 次
- ScaleLong: Towards More Stable Training of Diffusion Model via Scaling Network Long Skip ConnectionZhongzhan Huang, Pan Zhou, Shuicheng Yan, Liang LinNeurIPS 2023 · 被引用 41 次
- Understanding Self-attention Mechanism via Dynamical System PerspectiveZhongzhan Huang, Mingfu Liang, Jinghui Qin, Shanshan Zhong 等ICCV 2023 · 被引用 38 次
- Mirror Gradient: Towards Robust Multimodal Recommender Systems via Exploring Flat Local MinimaShanshan Zhong, Zhongzhan Huang, Daifeng Li, Wushao Wen 等WWW 2024 · 被引用 24 次
相关 Paper
- Cross-Layer Retrospective Retrieving via Layer AttentionYanwen Fang, Yuxi Cai, Jintai Chen, Jingyu Zhao 等ICLR 2023 · 被引用 1 次
- Recurrence along Depth: Deep Convolutional Neural Networks with Recurrent Layer AggregationJingyu Zhao, Yanwen Fang, Guodong LiNeurIPS 2021 · 被引用 31 次
- Loss-Based Attention for Deep Multiple Instance LearningXiaoshuang Shi, Fuyong Xing, Yuanpu Xie, Zizhao Zhang 等AAAI 2020 · 被引用 123 次
- AA-RMVSNet: Adaptive Aggregation Recurrent Multi-view Stereo NetworkZizhuang Wei, Qingtian Zhu, Chen Min, Yisong Chen 等ICCV 2021 · 被引用 193 次
- AttentionRNN: A Structured Spatial Attention MechanismSiddhesh Khandelwal, Leonid SigalICCV 2019 · 被引用 4 次
