Attentive Pooling with Learnable Norms for Text Representation
Chuhan Wu, Fangzhao Wu, Tao Qi, Xiaohui Cui, Yongfeng Huang
摘要
Pooling is an important technique for learning text representations in many neural NLP models. In conventional pooling methods such as average, max and attentive pooling, text representations are weighted summations of the L 1 or L ∞ norm of input features. However, their pooling norms are always fixed and may not be optimal for learning accurate text representations in different tasks. In addition, in many popular pooling methods such as max and attentive pooling some features may be over-emphasized, while other useful ones are not fully exploited. In this paper, we propose an Attentive Pooling with Learnable Norms (APLN) approach for text representation. Different from existing pooling methods that use a fixed pooling norm, we propose to learn the norm in an end-to-end manner to automatically find the optimal ones for text representation in different tasks. In addition, we propose two methods to ensure the numerical stability of the model training. The first one is scale limiting, which re-scales the input to ensure non-negativity and alleviate the risk of exponential explosion. The second one is re-formulation, which decomposes the exponent operation to avoid computing the realvalued powers of the input and further accelerate the pooling operation. Experimental results on four benchmark datasets show that our approach can effectively improve the performance of attentive pooling.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- HOP to the Next Tasks and Domains for Continual Learning in NLPUmberto Michieli, Mete OzayAAAI 2024 · 被引用 3 次
- Towards Efficient Dialogue Pre-training with Transferable and Interpretable Latent StructureXueliang Zhao, Lemao Liu, Tingchen Fu, Shuming Shi 等EMNLP 2022 · 被引用 3 次
相关 Paper
- Affine-Scaled Attention: Towards Flexible and Stable Transformer AttentionJeongin Bae, baeseong park, Gunho Park, Minsub Kim 等ICML 2026 · 被引用 1 次
- Efficient Representation Learning via Adaptive Context PoolingChen Huang, Walter Talbott, Navdeep Jaitly, Joshua M. SusskindICML 2022 · 被引用 10 次
- Pool Me Wisely: On the Effect of Pooling in Transformer-Based ModelsSofiane Ennadir, Levente Zólyomi, Oleg Smirnov, Tianze Wang 等NeurIPS 2025 · 被引用 6 次
- Towards Improved Sentence Representations using Token GraphsKrishna Sri Ipsit Mantri, Carola-Bibiane Schönlieb, Zorah Lähner, Moshe EliasofICLR 2026 · 被引用 1 次
- Why Mean Pooling Works: Quantifying Second-Order Collapse in Text EmbeddingsTomomasa Hara, Hiroto Kurita, Masaaki Imaizumi, Kentaro Inui 等ACL 2026 · 被引用 2 次
