VISTA: Verification In Sequential Turn-based Assessment
Ashley Lewis, Andrew Perrault, Eric Fosler-Lussier, Michael White
Abstract
Hallucination-defined here as generated statements unsupported or contradicted by available evidence or conversational context-remains a major obstacle to using conversational AI systems in settings that demand factual reliability. Existing metrics evaluate isolated responses or treat unverifiable content as errors, limiting their use for multi-turn dialogue. We introduce VISTA (Verification In Sequential Turn-based Assessment), a framework for evaluating conversational factuality via claim-level verification and sequential consistency tracking. VISTA decomposes each turn into atomic claims, verifies them against trusted sources and dialogue history, and categorizes unverifiable statements (subjective, contradicted, lacking evidence, or abstaining). Across eight large language models and four dialogue factuality benchmarks (AIS, BEGIN, FAITHDIAL, and FADE), VISTA substantially improves hallucination detection over FActScore and LLM-as-Judge baselines. Human evaluation confirms that VISTA's decomposition improves annotator agreement and reveals inconsistencies in existing benchmarks. Further analyses show that incorporating dialogue context into verification substantially improves contradiction detection, and that VISTA reliably identifies abstentions. By modeling factuality as a dynamic property of conversation, VISTA offers a more transparent, human-aligned measure of truthfulness in dialogue systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2c5daa4f-9c0e-43b1-947e-30494e14bc4dBuilds on8
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang et al.EMNLP 2023 · 549 citations
- SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language ModelsPotsawee Manakul, Adian Liusie, Mark J. F. GalesEMNLP 2023 · 331 citations
- FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text GenerationSewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis et al.EMNLP 2023 · 225 citations
- HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language ModelsJunyi Li, Xiaoxue Cheng, Xin Zhao, Jian-Yun Nie et al.EMNLP 2023 · 224 citations
- : Evaluating Factual Consistency in Knowledge-Grounded Dialogues via Question Generation and Question AnsweringOr Honovich, Leshem Choshen, Roee Aharoni, Ella Neeman et al.EMNLP 2021 · 101 citations
Related papers
- INFACT: A Diagnostic Benchmark for Induced Faithfulness and Factuality Hallucinations in Video-LLMsJunqi Yang, Yuecong Min, Jie Zhang, Shiguang Shan et al.ACL 2026 · 7 citations
- FactBench: A Dynamic Benchmark for In-the-Wild Language Model Factuality EvaluationFarima Fatahi Bayat, Lechen Zhang, Sheza Munir, Lu WangACL 2025
- Beyond Facts: Evaluating Intent Hallucination in Large Language ModelsYijie Hao, Haofei Yu, Jiaxuan YouACL 2025
- VisDiaHalBench: A Visual Dialogue Benchmark For Diagnosing Hallucination in Large Vision-Language ModelsQingxing Cao, Junhao Cheng, Xiaodan Liang, Liang LinACL 2024 · 3 citations
- LM vs LM: Detecting Factual Errors via Cross ExaminationRoi Cohen, May Hamri, Mor Geva, Amir GlobersonEMNLP 2023 · 41 citations
