TrueTeacher: Learning Factual Consistency Evaluation with Large Language Models
Zorik Gekhman, Jonathan Herzig, Roee Aharoni, Chen Elkind, Idan Szpektor
摘要
Factual consistency evaluation is often conducted using Natural Language Inference (NLI) models, yet these models exhibit limited success in evaluating summaries. Previous work improved such models with synthetic training data. However, the data is typically based on perturbed human-written summaries, which often differ in their characteristics from real model-generated summaries and have limited coverage of possible factual errors. Alternatively, large language models (LLMs) have recently shown promising results in directly evaluating generative tasks, but are too computationally expensive for practical use. Motivated by these limitations, we introduce TrueTeacher, a method for generating synthetic data by annotating diverse model-generated summaries using a LLM. Unlike prior work, TrueTeacher does not rely on human-written summaries, and is multilingual by nature. Experiments on the TRUE benchmark show that a student model trained using our data, substantially outperforms both the state-of-the-art model with similar capacity, and the LLM teacher. In a systematic study, we compare TrueTeacher to existing synthetic data generation methods and demonstrate its superiority and robustness to domain-shift. We also show that our method generalizes to multilingual scenarios. Lastly, we release our largescale synthetic dataset (1.4M examples), generated using TrueTeacher, and a checkpoint trained on this data. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper32
- Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?Zorik Gekhman, Gal Yona, Roee Aharoni, Matan Eyal 等EMNLP 2024 · 被引用 53 次
- IRCAN: Mitigating Knowledge Conflicts in LLM Generation via Identifying and Reweighting Context-Aware NeuronsDan Shi, Renren Jin, Tianhao Shen, Weilong Dong 等NeurIPS 2024 · 被引用 44 次
- LM vs LM: Detecting Factual Errors via Cross ExaminationRoi Cohen, May Hamri, Mor Geva, Amir GlobersonEMNLP 2023 · 被引用 41 次
- Knowledge Conflicts for LLMs: A SurveyRongwu Xu, Zehan Qi, Zhijiang Guo, Cunxiang Wang 等EMNLP 2024 · 被引用 38 次
- From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judgeDawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi 等EMNLP 2025 · 被引用 37 次
它引用的顶会 Paper13
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal 等ACL 2020 · 被引用 602 次
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang 等EMNLP 2023 · 被引用 549 次
相关 Paper
- Truth Knows No Language: Evaluating Truthfulness Beyond EnglishBlanca Calvo Figueras, Eneko Sagarzazu, Julen Etxaniz, Jeremy Barnes 等ACL 2025
- Are LLMs Better than Reported? Detecting Label Errors and Mitigating Their Effect on Model PerformanceOmer Nahum, Nitay Calderon, Orgad Keller, Idan Szpektor 等EMNLP 2025 · 被引用 9 次
- SummEdits: Measuring LLM Ability at Factual Reasoning Through The Lens of SummarizationPhilippe Laban, Wojciech Kryscinski, Divyansh Agarwal, Alexander R. Fabbri 等EMNLP 2023 · 被引用 29 次
- Autonomous Evaluation of LLMs for Truth Maintenance and Reasoning TasksRushang Karia, Daniel Bramblett, Daksh Dobhal, Siddharth SrivastavaICLR 2025
- WeCheck: Strong Factual Consistency Checker via Weakly Supervised LearningWenhao Wu, Wei Li, Xinyan Xiao, Jiachen Liu 等ACL 2023 · 被引用 3 次
