Improving Temporal Generalization of Pre-trained Language Models with Lexical Semantic Change
Zhaochen Su, Zecheng Tang, Xinyan Guan, Lijun Wu, Min Zhang, Juntao Li
摘要
Recent research has revealed that neural language models at scale suffer from poor temporal generalization capability, i.e., language model pre-trained on static data from past years performs worse over time on emerging data. Existing methods mainly perform continual training to mitigate such a misalignment. While effective to some extent but is far from being addressed on both the language modeling and downstream tasks. In this paper, we empirically observe that temporal generalization is closely affiliated with lexical semantic change, which is one of the essential phenomena of natural languages. Based on this observation, we propose a simple yet effective lexical-level masking strategy to post-train a converged language model. Experiments on two pre-trained language models, two different classification tasks, and four benchmark datasets demonstrate the effectiveness of our proposed method over existing temporal adaptation methods, i.e., continual training with new data. Our code is available at https: //github.com/zhaochen0110/LMLM .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Not All Contexts Are Equal: Teaching LLMs Credibility-aware GenerationRuotong Pan, Boxi Cao, Hongyu Lin, Xianpei Han 等EMNLP 2024 · 被引用 4 次
- SURf: Teaching Large Vision-Language Models to Selectively Utilize Retrieved InformationJiashuo Sun, Jihai Zhang, Yucheng Zhou, Zhaochen Su 等EMNLP 2024 · 被引用 2 次
- REAL: Resolving Knowledge Conflicts in Knowledge-Intensive Visual Question Answering via Reasoning-Pivot AlignmentKai Ye, Xianwei Mao, Sheng Zhou, Zirui Shao 等ICML 2026 · 被引用 1 次
- Knowledge-driven Augmentation and Retrieval for Integrative Temporal AdaptationWeisi Liu, Guangzeng Han, Xiaolei HuangACL 2026 · 被引用 1 次
- Once Upon a Time in Graph: Relative-Time Pretraining for Complex Temporal ReasoningSen Yang, Xin Li, Lidong Bing, Wai LamEMNLP 2023 · 被引用 1 次
它引用的顶会 Paper7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- Mind the Gap: Assessing Temporal Generalization in Neural Language ModelsAngeliki Lazaridou, Adhiguna Kuncoro, Elena Gribovskaya, Devang Agrawal 等NeurIPS 2021 · 被引用 315 次
- Analysing Lexical Semantic Change with Contextualised Word RepresentationsMario Giulianelli, Marco Del Tredici, Raquel FernándezACL 2020 · 被引用 118 次
- Temporally-Informed Analysis of Named Entity RecognitionShruti Rijhwani, Daniel Preotiuc-PietroACL 2020 · 被引用 49 次
相关 Paper
- Continual Pre-training of Language ModelsZixuan Ke, Yijia Shao, Haowei Lin, Tatsuya Konishi 等ICLR 2023 · 被引用 15 次
- Analyzing Semantic Change through Lexical ReplacementsFrancesco Periti, Pierluigi Cassotti, Haim Dubossarsky, Nina TahmasebiACL 2024
- NEO-BENCH: Evaluating Robustness of Large Language Models with NeologismsJonathan Zheng, Alan Ritter, Wei XuACL 2024
- On-the-fly Cross-lingual Masking for Multilingual Pre-trainingXi Ai, Bin FangACL 2023 · 被引用 1 次
- MASKER: Masked Keyword Regularization for Reliable Text ClassificationSeung Jun Moon, Sangwoo Mo, Kimin Lee, Jaeho Lee 等AAAI 2021 · 被引用 39 次
