Factual or Contextual? Disentangling Error Types in Entity Description Generation
Navita Goyal, Ani Nenkova, Hal Daumé III
摘要
In the task of entity description generation, given a context and a specified entity, a model must describe that entity correctly and in a contextually-relevant way. In this task, as well as broader language generation tasks, the generation of a nonfactual description (factual error) versus an incongruous description (contextual error) is fundamentally different, yet often conflated. We develop an evaluation paradigm that enables us to disentangle these two types of errors in naturally occurring textual contexts. We find that factuality and congruity are often at odds, and that models specifically struggle with accurate descriptions of entities that are less familiar to people. This shortcoming of language models raises concerns around the trustworthiness of such models, since factual errors on less well-known entities are exactly those that a human reader will not recognize. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Tool Unlearning for Tool-Augmented LLMsJiali Cheng, Hadi AmiriICML 2025
- Intrinsic Task-based Evaluation for Referring Expression GenerationGuanyi Chen, Fahime Same, Kees van DeemterACL 2024
它引用的顶会 Paper6
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Large Language Models Struggle to Learn Long-Tail KnowledgeNikhil Kandpal, Haikang Deng, Adam Roberts, Eric Wallace 等ICML 2023 · 被引用 623 次
- Neural Path Hunter: Reducing Hallucination in Dialogue Systems via Path GroundingNouha Dziri, Andrea Madotto, Osmar Zaïane, Avishek Joey BoseEMNLP 2021 · 被引用 74 次
- On Faithfulness and Factuality in Abstractive SummarizationJoshua Maynez, Shashi Narayan, Bernd Bohnet, Ryan T. McDonaldACL 2020 · 被引用 54 次
相关 Paper
- DisentQA: Disentangling Parametric and Contextual Knowledge with Counterfactual Question AnsweringElla Neeman, Roee Aharoni, Or Honovich, Leshem Choshen 等ACL 2023 · 被引用 20 次
- TrueTeacher: Learning Factual Consistency Evaluation with Large Language ModelsZorik Gekhman, Jonathan Herzig, Roee Aharoni, Chen Elkind 等EMNLP 2023 · 被引用 24 次
- Models See Hallucinations: Evaluating the Factuality in Video CaptioningHui Liu, Xiaojun WanEMNLP 2023 · 被引用 5 次
- Beyond Factuality: A Comprehensive Evaluation of Large Language Models as Knowledge GeneratorsLiang Chen, Yang Deng, Yatao Bian, Zeyu Qin 等EMNLP 2023 · 被引用 21 次
- FaithEval: Can Your Language Model Stay Faithful to Context, Even If "The Moon is Made of Marshmallows"Yifei Ming, Senthil Purushwalkam, Shrey Pandit, Zixuan Ke 等ICLR 2025
