Will It Still Be True Tomorrow? Multilingual Evergreen Question Classification to Improve Trustworthy QA
Sergey Pletenev, Maria Marina, Nikolay Ivanov, Daria Galimzianova, Nikita Krayko, Mikhail Salnikov, Vasily Konovalov, Alexander Panchenko, Viktor Moskvoretskii
Abstract
Large Language Models (LLMs) often hallucinate in question answering (QA) tasks. A key yet underexplored factor contributing to this is the temporality of questions -whether they are evergreen (answers remain stable over time) or mutable (answers change). In this work, we introduce EverGreenQA, the first multilingual QA dataset with evergreen labels, supporting both evaluation and training. Using Ever-GreenQA, we benchmark 12 modern LLMs to assess whether they encode question temporality explicitly (via verbalized judgments) or implicitly (via uncertainty signals). We also train EG-E5, a lightweight multilingual classifier that achieves SoTA performance on this task. Finally, we demonstrate the practical utility of evergreen classification across three applications: improving self-knowledge estimation, filtering QA datasets, and explaining GPT-4o's retrieval behavior. * Equal contribution. ** Work has been done while at Skoltech.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on7
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 394 citations
- Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step QuestionsHarsh Trivedi, Niranjan Balasubramanian, Tushar Khot, Ashish SabharwalACL 2023 · 187 citations
- StreamingQA: A Benchmark for Adaptation to New Knowledge over Time in Question Answering ModelsAdam Liska, Tomás Kociský, Elena Gribovskaya, Tayfun Terzi et al.ICML 2022 · 129 citations
- Adaptive Retrieval Without Self-Knowledge? Bringing Uncertainty Back HomeViktor Moskvoretskii, Maria Marina, Mikhail Salnikov, Nikolay Ivanov et al.ACL 2025 · 22 citations
- SituatedQA: Incorporating Extra-Linguistic Contexts into QAMichael J. Q. Zhang, Eunsol ChoiEMNLP 2021 · 2 citations
Related papers
- How Much Do LLMs Hallucinate across Languages? On Realistic Multilingual Estimation of LLM HallucinationSaad Obaid ul Islam, Anne Lauscher, Goran GlavasEMNLP 2025
- ANAH: Analytical Annotation of Hallucinations in Large Language ModelsZiwei Ji, Yuzhe Gu, Wenwei Zhang, Chengqi Lyu et al.ACL 2024 · 8 citations
- K-HALU: Multiple Answer Korean Hallucination Benchmark for Large Language ModelsJaehyung Seo, Heuiseok LimICLR 2025
- XLQA: A Benchmark for Locale-Aware Multilingual Open-Domain Question AnsweringKeon-Woo Roh, Yeong-Joon Ju, Seong-Whan LeeEMNLP 2025
- Beyond Facts: Evaluating Intent Hallucination in Large Language ModelsYijie Hao, Haofei Yu, Jiaxuan YouACL 2025
