Enhancing Elusive Clues in Knowledge Learning by Contrasting Attention of Language Models
Jian Gao, Xiao Zhang, Miao Li, Ji Wu
摘要
Causal language models acquire vast amount of knowledge from general text corpus during pretraining, but the efficiency of knowledge learning is known to be unsatisfactory, especially when learning from knowledge-dense and small-sized corpora. The deficiency can come from long-distance dependencies which are hard to capture by language models, and overfitting to co-occurrence patterns and distracting clues in the training text. To address these issues, the paper proposes a method to enhance knowledge learning during language model pretraining, by enhancing elusive but important clues in text discovered by the language model themselves. We found that larger language models pay more attention to non-obvious but important clues, which are often overlooked by smaller language models. Therefore, we can identify these clues by contrasting the attention weights of large and small language models. We use the identified clues as a guide to perform token-dropout data augmentation on the training text, and observed a significant boost in both small and large models' performance in fact memorization. This shows that the behavior contrast between more and less-performant language models contains important clues for knowledge learning, and it can be "amplified" for a straight-forward improvement in knowledge learning efficiency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong 等NeurIPS 2020 · 被引用 2,774 次
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer 等NeurIPS 2022 · 被引用 2,039 次
- MAmmoTH: Building Math Generalist Models through Hybrid Instruction TuningXiang Yue, Xingwei Qu, Ge Zhang, Yao Fu 等ICLR 2024 · 被引用 558 次
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu 等ACL 2023 · 被引用 540 次
相关 Paper
- Analysing The Impact of Sequence Composition on Language Model Pre-TrainingYu Zhao, Yuanbin Qu, Konrad Staniszewski, Szymon Tworkowski 等ACL 2024
- How Do Large Language Models Acquire Factual Knowledge During Pretraining?Hoyeon Chang, Jinho Park, Seonghyeon Ye, Sohee Yang 等NeurIPS 2024 · 被引用 124 次
- Large Language Models Struggle to Learn Long-Tail KnowledgeNikhil Kandpal, Haikang Deng, Adam Roberts, Eric Wallace 等ICML 2023 · 被引用 623 次
- Synthetic continued pretrainingZitong Yang, Neil Band, Shuangping Li, Emmanuel J. Candès 等ICLR 2025
- Pre-training Language Models with Deterministic Factual KnowledgeShaobo Li, Xiaoguang Li, Lifeng Shang, Chengjie Sun 等EMNLP 2022 · 被引用 13 次
