Robust Noise Attenuation via Adaptive Pooling of Transformer Outputs
Greyson Brothers
Abstract
We investigate the design of pooling methods used to summarize the outputs of transformer embedding models, primarily motivated by reinforcement learning and vision applications. This work considers problems where a subset of the input vectors contains requisite information for a downstream task (signal) while the rest are distractors (noise). By framing pooling as vector quantization with the goal of minimizing signal loss, we demonstrate that the standard methods used to aggregate transformer outputs, AvgPool, MaxPool, and ClsToken, are vulnerable to performance collapse as the signal-to-noise ratio (SNR) of inputs fluctuates. We then show that an attention-based adaptive pooling method can approximate the signal-optimal vector quantizer within derived error bounds for any SNR. Our theoretical results are first validated by supervised experiments on a synthetic dataset designed to isolate the SNR problem, then generalized to standard relational reasoning, multi-agent reinforcement learning, and vision benchmarks with noisy observations, where transformers with adaptive pooling display superior robustness across tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cbfc6c9d-0aba-433b-9aa4-2166c1985378Cited by top-tier papers2
- AdaJudge: Adaptive Multi-Perspective Judging for Reward ModelingYongliang Miao, Yangyang Liang, Mengnan DuACL 2026 · 1 citation
- Towards Improved Sentence Representations using Token GraphsKrishna Sri Ipsit Mantri, Carola-Bibiane Schönlieb, Zorah Lähner, Moshe EliasofICLR 2026 · 1 citation
Builds on12
- Going deeper with Image TransformersHugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve et al.ICCV 2021 · 1,279 citations
- Scaling Vision TransformersXiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas BeyerCVPR 2022 · 767 citations
- Hopfield Networks is All You NeedHubert Ramsauer, Bernhard Schäfl, Johannes Lehner, Philipp Seidl et al.ICLR 2021 · 620 citations
- GameFormer: Game-theoretic Modeling and Learning of Transformer-based Interactive Prediction and Planning for Autonomous DrivingZhiyu Huang, Haochen Liu, Chen LvICCV 2023 · 209 citations
- Scalable Multi-Agent Reinforcement Learning through Intelligent Information AggregationSiddharth Nayak, Kenneth Choi, Wenqi Ding, Sydney Dolan et al.ICML 2023 · 73 citations
Related papers
- Pool Me Wisely: On the Effect of Pooling in Transformer-Based ModelsSofiane Ennadir, Levente Zólyomi, Oleg Smirnov, Tianze Wang et al.NeurIPS 2025 · 6 citations
- Why Mean Pooling Works: Quantifying Second-Order Collapse in Text EmbeddingsTomomasa Hara, Hiroto Kurita, Masaaki Imaizumi, Kentaro Inui et al.ACL 2026 · 2 citations
- Unlocking Noise-Resistant Vision: Key Architectural Secrets for Robust Models Against Gaussian NoiseBum Jun Kim, Makoto Kawano, Yusuke Iwasawa, Yutaka MatsuoICML 2026
- Keep It SimPool: Who Said Supervised Transformers Suffer from Attention Deficit?Bill Psomas, Ioannis Kakogeorgiou, Konstantinos Karantzalos, Yannis AvrithisICCV 2023 · 18 citations
- Robustifying Token Attention for Vision TransformersYong Guo, David Stutz, Bernt SchieleICCV 2023 · 37 citations
