PACEAttention: Principled and Adaptive Feature Compression-Expansion Grounded in the Geometry of
Xiaojie Yu, Haibo Zhang, Jeremiah D. Deng, Lizhi Peng
摘要
The maximal coding rate reduction (MCR 2 ) objective is proposed for learning low-dimensional subspace representations and for principled deep model design, where layer structures are derived by unrolling its optimization steps. However, existing methods motivated by this objective do not fully adhere to design principles implied by the MCR 2 gradient, which weakens the principled and interpretable foundations of the resulting models. In this work, we introduce PACEAttention, a novel principled attention mechanism inspired by the geometric insight of MCR 2 , whose gradientbased updates move features along directions shaped by the underlying low-dimensional feature structure. Our method captures this structure by leveraging randomization to guide feature updates. This principled construction enables the resulting PACENet to exhibit enhanced interpretability, with different heads attending to distinct image regions and capturing fine-grained structures under simple supervised training. Experiments demonstrate that PACEAttention achieves superior performance and more stable scalability than previous principled modules while remaining low complexity. Code is available at this https URL.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 被引用 2,665 次
- XCiT: Cross-Covariance Image TransformersAlaaeldin Ali, Hugo Touvron, Mathilde Caron, Piotr Bojanowski 等NeurIPS 2021 · 被引用 692 次
- The Intrinsic Dimension of Images and Its Impact on LearningPhillip Pope, Chen Zhu, Ahmed Abdelkader, Micah Goldblum 等ICLR 2021 · 被引用 381 次
- White-Box Transformers via Sparse Rate ReductionYaodong Yu, Sam Buchanan, Druv Pai, Tianzhe Chu 等NeurIPS 2023 · 被引用 149 次
- The Tunnel Effect: Building Data Representations in Deep Neural NetworksWojciech Masarczyk, Mateusz Ostaszewski, Ehsan Imani, Razvan Pascanu 等NeurIPS 2023 · 被引用 40 次
相关 Paper
- Multi-ReduNet: Interpretable Class-Wise Decomposition of ReduNetFengrong Li, Delin ChuICLR 2026
- A Global Geometric Analysis of Maximal Coding Rate ReductionPeng Wang, Huikang Liu, Druv Pai, Yaodong Yu 等ICML 2024 · 被引用 13 次
- Efficient Maximal Coding Rate Reduction by Variational FormsChristina Baek, Ziyang Wu, Kwan Ho Ryan Chan, Tianjiao Ding 等CVPR 2022 · 被引用 5 次
- Rethinking Decoders for Transformer-based Semantic Segmentation: A Compression PerspectiveQishuai Wen, Chun-Guang LiNeurIPS 2024 · 被引用 8 次
- Learning Diverse and Discriminative Representations via the Principle of Maximal Coding Rate ReductionYaodong Yu, Kwan Ho Ryan Chan, Chong You, Chaobing Song 等NeurIPS 2020 · 被引用 265 次
