Multi-Level Head-Wise Match and Aggregation in Transformer for Textual Sequence Matching
Shuohang Wang, Yunshi Lan, Yi Tay, Jing Jiang, Jingjing Liu
Abstract
Transformer has been successfully applied to many natural language processing tasks. However, for textual sequence matching, simple matching between the representation of a pair of sequences might bring in unnecessary noise. In this paper, we propose a new approach to sequence pair matching with Transformer, by learning head-wise matching representations on multiple levels. Experiments show that our proposed approach can achieve new state-of-the-art performance on multiple tasks that rely only on pre-computed sequence-vector-representation, such as SNLI, MNLI-match, MNLI-mismatch, QQP, and SQuAD-binary.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 35819a32-415d-4ba2-8ed4-269c2bf8b793Related papers
- Poly-encoders: Architectures and Pre-training Strategies for Fast and Accurate Multi-sentence ScoringSamuel Humeau, Kurt Shuster, Marie-Anne Lachaux, Jason WestonICLR 2020 · 316 citations
- Layered Insights: Generalizable Analysis of Human Authorial Style by Leveraging All Transformer LayersMilad Alshomary, Nikhil Reddy Varimalla, Vishal Anand, Smaranda Muresan et al.EMNLP 2025
- Using Prior Knowledge to Guide BERT's Attention in Semantic Textual Matching TasksTingyu Xia, Yue Wang, Yuan Tian, Yi ChangWWW 2021 · 56 citations
- Paraphrase Generation by Learning How to Edit from SamplesAmirhossein Kazemnejad, Mohammadreza Salehi, Mahdieh Soleymani BaghshahACL 2020 · 43 citations
- Cross-Thought for Sentence Encoder Pre-trainingShuohang Wang, Yuwei Fang, Siqi Sun, Zhe Gan et al.EMNLP 2020 · 17 citations
