ACT: an Attentive Convolutional Transformer for Efficient Text Classification
Pengfei Li, Peixiang Zhong, Kezhi Mao, Dongzhe Wang, Xuefeng Yang, Yunfeng Liu, Jianxiong Yin, Simon See
摘要
Recently, Transformer has been demonstrating promising performance in many NLP tasks and showing a trend of replacing Recurrent Neural Network (RNN). Meanwhile, less attention is drawn to Convolutional Neural Network (CNN) due to its weak ability in capturing sequential and long-distance dependencies, although it has excellent local feature extraction capability. In this paper, we introduce an Attentive Convolutional Transformer (ACT) that takes the advantages of both Transformer and CNN for efficient text classification. Specifically, we propose a novel attentive convolution mechanism that utilizes the semantic meaning of convolutional filters attentively to transform text from complex word space to a more informative convolutional filter space where important n-grams are captured. ACT is able to capture both local and global dependencies effectively while preserving sequential information. Experiments on various text classification tasks and detailed analyses show that ACT is a lightweight, fast, and effective universal text classifier, outperforming CNNs, RNNs, and attentive models including Transformer.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Transferable Post-hoc Calibration on Pretrained Transformers in Noisy Text ClassificationJun Zhang, Wen Yao, Xiaoqian Chen, Ling FengAAAI 2023 · 被引用 5 次
- Inceptive Transformers: Enhancing Contextual Representations through Multi-Scale Feature Learning Across Domains and LanguagesAsif Shahriar, Rifat Shahriyar, M. Saifur RahmanEMNLP 2025
相关 Paper
- Not All Attention Is Needed: Gated Attention Network for Sequence DataLanqing Xue, Xiaopeng Li, Nevin L. ZhangAAAI 2020 · 被引用 47 次
- GATE: Graph Attention Transformer Encoder for Cross-lingual Relation and Event ExtractionWasi Uddin Ahmad, Nanyun Peng, Kai-Wei ChangAAAI 2021 · 被引用 113 次
- Bag-of-Words vs. Graph vs. Sequence in Text Classification: Questioning the Necessity of Text-Graphs and the Surprising Strength of a Wide MLPLukas Galke, Ansgar ScherpACL 2022
- Convolutions and Self-Attention: Re-interpreting Relative Positions in Pre-trained Language ModelsTyler A. Chang, Yifan Xu, Weijian Xu, Zhuowen TuACL 2021
- GTC: Guided Training of CTC towards Efficient and Accurate Scene Text RecognitionWenyang Hu, Xiaocong Cai, Jun Hou, Shuai Yi 等AAAI 2020 · 被引用 151 次
