Using Prior Knowledge to Guide BERT's Attention in Semantic Textual Matching Tasks
Tingyu Xia, Yue Wang, Yuan Tian, Yi Chang
摘要
We study the problem of incorporating prior knowledge into a deep Transformer-based model, i.e., Bidirectional Encoder Representations from Transformers (BERT), to enhance its performance on semantic textual matching tasks. By probing and analyzing what BERT has already known when solving this task, we obtain better understanding of what task-specific knowledge BERT needs the most and where it is most needed. The analysis further motivates us to take a different approach than most existing works. Instead of using prior knowledge to create a new training task for fine-tuning BERT, we directly inject knowledge into BERT’s multi-head attention mechanism. This leads us to a simple yet effective approach that enjoys fast training stage as it saves the model from training on additional data or tasks other than the main task. Extensive experiments demonstrate that the proposed knowledge-enhanced BERT is able to consistently improve semantic textual matching performance over the original BERT model, and the performance benefit is most salient when training data is scarce.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Dict-TTS: Learning to Pronounce with Prior Dictionary Knowledge for Text-to-SpeechZiyue Jiang, Su Zhe, Zhou Zhao, Qian Yang 等NeurIPS 2022 · 被引用 7 次
- Learning Semantic Textual Similarity via Topic-informed Discrete Latent VariablesErxin Yu, Lan Du, Yuan Jin, Zhepei Wei 等EMNLP 2022 · 被引用 5 次
- Predicate-Argument Based Bi-Encoder for Paraphrase IdentificationQiwei Peng, David J. Weir, Julie Weeds, Yekun ChaiACL 2022
- Embedding Prior Task-specific Knowledge into Language Models for Context-aware Document RankingShuting Wang, Yutao Zhu, Zhicheng DouKDD 2025
- TestifAI: Tomography-Based Testing for Deep Learning SystemsArooj Arif, Tobias Hartung, Elena Botoeva, Alexandros KoliousisICSE 2026
它引用的顶会 Paper6
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 被引用 3,729 次
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong 等NeurIPS 2020 · 被引用 2,774 次
- K-BERT: Enabling Language Representation with Knowledge GraphWeijie Liu, Peng Zhou, Zhe Zhao, Zhiruo Wang 等AAAI 2020 · 被引用 898 次
- Semantics-Aware BERT for Language UnderstandingZhuosheng Zhang, Yuwei Wu, Hai Zhao, Zuchao Li 等AAAI 2020 · 被引用 396 次
相关 Paper
- Analyzing How BERT Performs Entity MatchingMatteo Paganelli, Francesco Del Buono, Andrea Baraldi, Francesco GuerraVLDB 2022 · 被引用 35 次
- Deep Entity Matching with Pre-Trained Language ModelsYuliang Li, Jinfeng Li, Yoshihiko Suhara, AnHai Doan 等VLDB 2021 · 被引用 484 次
- Enhancing Transformer-based Semantic Matching for Few-shot Learning through Weakly Contrastive Pre-trainingWei Yang, Tengfei Huo, Zhiqiang LiuACM MM 2024
- Acquiring Knowledge from Pre-Trained Model to Neural Machine TranslationRongxiang Weng, Heng Yu, Shujian Huang, Shanbo Cheng 等AAAI 2020 · 被引用 71 次
- Incorporating medical knowledge in BERT for clinical relation extractionArpita Roy, Shimei PanEMNLP 2021 · 被引用 56 次
