Improving Temporal Generalization of Pre-trained Language Models with Lexical Semantic Change
Zhaochen Su, Zecheng Tang, Xinyan Guan, Lijun Wu, Min Zhang, Juntao Li
Abstract
Recent research has revealed that neural language models at scale suffer from poor temporal generalization capability, i.e., language model pre-trained on static data from past years performs worse over time on emerging data. Existing methods mainly perform continual training to mitigate such a misalignment. While effective to some extent but is far from being addressed on both the language modeling and downstream tasks. In this paper, we empirically observe that temporal generalization is closely affiliated with lexical semantic change, which is one of the essential phenomena of natural languages. Based on this observation, we propose a simple yet effective lexical-level masking strategy to post-train a converged language model. Experiments on two pre-trained language models, two different classification tasks, and four benchmark datasets demonstrate the effectiveness of our proposed method over existing temporal adaptation methods, i.e., continual training with new data. Our code is available at https: //github.com/zhaochen0110/LMLM .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Not All Contexts Are Equal: Teaching LLMs Credibility-aware GenerationRuotong Pan, Boxi Cao, Hongyu Lin, Xianpei Han et al.EMNLP 2024 · 4 citations
- SURf: Teaching Large Vision-Language Models to Selectively Utilize Retrieved InformationJiashuo Sun, Jihai Zhang, Yucheng Zhou, Zhaochen Su et al.EMNLP 2024 · 2 citations
- REAL: Resolving Knowledge Conflicts in Knowledge-Intensive Visual Question Answering via Reasoning-Pivot AlignmentKai Ye, Xianwei Mao, Sheng Zhou, Zirui Shao et al.ICML 2026 · 1 citation
- Knowledge-driven Augmentation and Retrieval for Integrative Temporal AdaptationWeisi Liu, Guangzeng Han, Xiaolei HuangACL 2026 · 1 citation
- Once Upon a Time in Graph: Relative-Time Pretraining for Complex Temporal ReasoningSen Yang, Xin Li, Lidong Bing, Wai LamEMNLP 2023 · 1 citation
Builds on7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Mind the Gap: Assessing Temporal Generalization in Neural Language ModelsAngeliki Lazaridou, Adhiguna Kuncoro, Elena Gribovskaya, Devang Agrawal et al.NeurIPS 2021 · 315 citations
- Analysing Lexical Semantic Change with Contextualised Word RepresentationsMario Giulianelli, Marco Del Tredici, Raquel FernándezACL 2020 · 118 citations
- Temporally-Informed Analysis of Named Entity RecognitionShruti Rijhwani, Daniel Preotiuc-PietroACL 2020 · 49 citations
Related papers
- Continual Pre-training of Language ModelsZixuan Ke, Yijia Shao, Haowei Lin, Tatsuya Konishi et al.ICLR 2023 · 15 citations
- Analyzing Semantic Change through Lexical ReplacementsFrancesco Periti, Pierluigi Cassotti, Haim Dubossarsky, Nina TahmasebiACL 2024
- NEO-BENCH: Evaluating Robustness of Large Language Models with NeologismsJonathan Zheng, Alan Ritter, Wei XuACL 2024
- On-the-fly Cross-lingual Masking for Multilingual Pre-trainingXi Ai, Bin FangACL 2023 · 1 citation
- MASKER: Masked Keyword Regularization for Reliable Text ClassificationSeung Jun Moon, Sangwoo Mo, Kimin Lee, Jaeho Lee et al.AAAI 2021 · 39 citations
