WikiBigEdit: Understanding the Limits of Lifelong Knowledge Editing in LLMs
Lukas Thede, Karsten Roth, Matthias Bethge, Zeynep Akata, Thomas Hartvigsen
Abstract
Keeping large language models factually up-todate is crucial for deployment, yet costly retraining remains a challenge. Knowledge editing offers a promising alternative, but methods are only tested on small-scale or synthetic edit benchmarks. In this work, we aim to bridge research into lifelong knowledge editing to real-world edits at a practically relevant scale. We first introduce WikiBigEdit; a large-scale benchmark of realworld Wikidata edits, built to automatically extend lifelong for future-proof benchmarking. In its first instance, it includes over 500K questionanswer pairs for knowledge editing alongside a comprehensive evaluation pipeline. Finally, we use WikiBigEdit to study existing knowledge editing techniques' ability to incorporate large volumes of real-world facts and contrast their capabilities to generic modification techniques such as retrieval augmentation and continual finetuning to acquire a complete picture of the practical extent of current lifelong knowledge editing. 1 Derived from periodic changes to Wikidata knowledge graphs (Jang et al., 2022; Khodja et al., 2024) , WikiBigEdit covers a large range of factual edits and refinements. Moreover, WikiBigEdit introduces comprehensive evaluation axes going beyond standard knowl-
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e73e3656-b727-4b7f-8139-fb807642cf74Cited by top-tier papers4
- Fine-tuning Done Right in Model EditingWanli Yang, Rui Tang, Hongyu Zang, Du Su et al.ICLR 2026 · 9 citations
- Aligning Language Models with Real-time Knowledge EditingChenming Tang, Yutong Yang, Kexue Wang, Yunfang WuACL 2026
- Representation Interventions Enable Lifelong Knowledge Memory Control in LLMsXuyuan Liu, Shengyu Chen, Xinshuai Dong, Yanchi Liu et al.ACL 2026
- CrispEdit: Low-Curvature Projections for Scalable Non-Destructive LLM EditingZarif Ikram, Arad Firouzkouhi, Stephen Tu, Mahdi Soltanolkotabi et al.ICML 2026
Builds on24
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat et al.ICML 2020 · 2,937 citations
- Improving Language Models by Retrieving from Trillions of TokensSebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai et al.ICML 2022 · 1,629 citations
Related papers
- Towards Scalable Lifelong Knowledge Editing with Selective Knowledge SuppressionDahyun Jung, Jaewook Lee, Heuiseok LimACL 2026
- MMKE-Bench: A Multimodal Editing Benchmark for Diverse Visual KnowledgeYuntao Du, Kailin Jiang, Zhi Gao, Chenrui Shi et al.ICLR 2025
- AKEW: Assessing Knowledge Editing in the WildXiaobao Wu, Liangming Pan, William Yang Wang, Anh Tuan LuuEMNLP 2024 · 2 citations
- Can Knowledge Editing Really Correct Hallucinations?Baixiang Huang, Canyu Chen, Xiongxiao Xu, Ali Payani et al.ICLR 2025
- MultiMedBench: A Scenario-Aware Benchmark for Evaluating Knowledge Editing in Medical VQAShengtao Wen, Haodong Chen, Yadong Wang, Zhongying Pan et al.AAAI 2026
