Selection-as-Nonlinearity: Bridging Attention and Activation via a Joint Game–Decision Lens for Interpretable, Discriminative Visual Representations
Sudong Cai, Shuai Yuan, Bingzhi Chen, Rui Mao, Bing Wang
Abstract
Self-attention with separate pre- and post-projections can be a universal approximator (on compact domains) under mild conditions.Yet we observe a striking gap: an attention-only Transformer (w/o FFN layers) exhibits a marked accuracy drop relative to its standard interleaved attention--FFN baseline.We term this the weak-independence challenge of attention.We study this through a new conceptual lens, Selection-as-Nonlinearity (SaN), which interprets effective nonlinearity as directed, cost-constrained selection, offering a coherent account of attention as context-gated activation.In this joint game–decision view, attention performs a resource-constrained cooperative allocation over values: each query distributes a unit-mass weight budget over shared values to optimize representational utility, under a normalizer (e.g., ), and guided by context-derived scores (e.g., q-k similarities).SaN interprets weak-independence as a structural tension: the value weights almost cannot simultaneously attain the decoupled per-query (row-wise) and the per-value (column-wise) optimums under shared budgets, thereby limiting attention's stand-alone capacity.Guided by SaN, we introduce CSaN, an interpretable, efficient attention compensation paradigm with two key insights: 1) hierarchical budget calibration, re-allocate row budgets via inter-query correction signals; and 2) public-private cooperation, enhancing the public attention pathway with a per-token private value pathway to decouple conflicting demands.CSaN is evaluated on various vision benchmarks and demonstrates level-jump gains across popular Transformer families (Swin, ViT, Hiera), enabling models to rival much heavier same-family counterparts as large.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on24
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
Related papers
- Selective Attention: Enhancing Transformer through Principled Context ControlXuechen Zhang, Xiangyu Chang, Mingchen Li, Amit K. Roy-Chowdhury et al.NeurIPS 2024 · 32 citations
- Scatterbrain: Unifying Sparse and Low-rank AttentionBeidi Chen, Tri Dao, Eric Winsor, Zhao Song et al.NeurIPS 2021 · 165 citations
- FasterViT: Fast Vision Transformers with Hierarchical AttentionAli Hatamizadeh, Greg Heinrich, Hongxu Yin, Andrew Tao et al.ICLR 2024 · 132 citations
- Cooperative or Competitive? Understanding the Interaction between Attention Heads From A Game Theory PerspectiveXiaoye Qu, Zengqi Yu, Dongrui Liu, Wei Wei et al.ACL 2025
- Learned Queries for Efficient Local AttentionMoab Arar, Ariel Shamir, Amit H. BermanoCVPR 2022 · 28 citations
