Using Prior Knowledge to Guide BERT's Attention in Semantic Textual Matching Tasks
Tingyu Xia, Yue Wang, Yuan Tian, Yi Chang
Abstract
We study the problem of incorporating prior knowledge into a deep Transformer-based model, i.e., Bidirectional Encoder Representations from Transformers (BERT), to enhance its performance on semantic textual matching tasks. By probing and analyzing what BERT has already known when solving this task, we obtain better understanding of what task-specific knowledge BERT needs the most and where it is most needed. The analysis further motivates us to take a different approach than most existing works. Instead of using prior knowledge to create a new training task for fine-tuning BERT, we directly inject knowledge into BERT’s multi-head attention mechanism. This leads us to a simple yet effective approach that enjoys fast training stage as it saves the model from training on additional data or tasks other than the main task. Extensive experiments demonstrate that the proposed knowledge-enhanced BERT is able to consistently improve semantic textual matching performance over the original BERT model, and the performance benefit is most salient when training data is scarce.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e228590e-e1b2-475a-b40f-4c4087358ca1Cited by top-tier papers5
- Dict-TTS: Learning to Pronounce with Prior Dictionary Knowledge for Text-to-SpeechZiyue Jiang, Su Zhe, Zhou Zhao, Qian Yang et al.NeurIPS 2022 · 7 citations
- Learning Semantic Textual Similarity via Topic-informed Discrete Latent VariablesErxin Yu, Lan Du, Yuan Jin, Zhepei Wei et al.EMNLP 2022 · 5 citations
- Predicate-Argument Based Bi-Encoder for Paraphrase IdentificationQiwei Peng, David J. Weir, Julie Weeds, Yekun ChaiACL 2022
- Embedding Prior Task-specific Knowledge into Language Models for Context-aware Document RankingShuting Wang, Yutao Zhu, Zhicheng DouKDD 2025
- TestifAI: Tomography-Based Testing for Deep Learning SystemsArooj Arif, Tobias Hartung, Elena Botoeva, Alexandros KoliousisICSE 2026
Builds on6
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong et al.NeurIPS 2020 · 2,774 citations
- K-BERT: Enabling Language Representation with Knowledge GraphWeijie Liu, Peng Zhou, Zhe Zhao, Zhiruo Wang et al.AAAI 2020 · 898 citations
- Semantics-Aware BERT for Language UnderstandingZhuosheng Zhang, Yuwei Wu, Hai Zhao, Zuchao Li et al.AAAI 2020 · 396 citations
Related papers
- Analyzing How BERT Performs Entity MatchingMatteo Paganelli, Francesco Del Buono, Andrea Baraldi, Francesco GuerraVLDB 2022 · 35 citations
- Deep Entity Matching with Pre-Trained Language ModelsYuliang Li, Jinfeng Li, Yoshihiko Suhara, AnHai Doan et al.VLDB 2021 · 484 citations
- Enhancing Transformer-based Semantic Matching for Few-shot Learning through Weakly Contrastive Pre-trainingWei Yang, Tengfei Huo, Zhiqiang LiuACM MM 2024
- Acquiring Knowledge from Pre-Trained Model to Neural Machine TranslationRongxiang Weng, Heng Yu, Shujian Huang, Shanbo Cheng et al.AAAI 2020 · 71 citations
- Incorporating medical knowledge in BERT for clinical relation extractionArpita Roy, Shimei PanEMNLP 2021 · 56 citations
