Information-Theoretic Text Hallucination Reduction for Video-grounded Dialogue
Sunjae Yoon, Eunseop Yoon, Hee Suk Yoon, Junyeong Kim, Chang Dong Yoo
Abstract
Video-grounded Dialogue (VGD) aims to decode an answer sentence to a question regarding a given video and dialogue context. Despite the recent success of multi-modal reasoning to generate answer sentences, existing dialogue systems still suffer from a text hallucination problem, which denotes indiscriminate text-copying from input texts without an understanding of the question. This is due to learning spurious correlations from the fact that answer sentences in the dataset usually include the words of input texts, thus the VGD system excessively relies on copying words from input texts by hoping those words to overlap with ground-truth texts. Hence, we design Text Hallucination Mitigating (THAM) framework, which incorporates Text Hallucination Regularization (THR) loss derived from the proposed information-theoretic text hallucination measurement approach. Applying THAM with current dialogue systems validates the effectiveness on VGD benchmarks (i.e., AVSD@DSTC7 and AVSD@DSTC8) and shows enhanced interpretability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- TPC: Test-time Procrustes Calibration for Diffusion-based Human Image AnimationSunjae Yoon, Gwanhyeong Koo, Younghwan Lee, Chang Dong YooNeurIPS 2024 · 16 citations
- FRAG: Frequency Adapting Group for Diffusion Video EditingSunjae Yoon, Gwanhyeong Koo, Geonwoo Kim, Chang D. YooICML 2024 · 13 citations
- An Audit on the Perspectives and Challenges of Hallucinations in NLPPranav Narayanan Venkit, Tatiana Chakravorti, Vipul Gupta, Heidi Biggs et al.EMNLP 2024 · 8 citations
- ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference OptimizationHee Suk Yoon, Eunseop Yoon, Mark A. Hasegawa-Johnson, Sungwoong Kim et al.ICML 2025
- Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language ModelsEunseop Yoon, Hee Suk Yoon, Mark A. Hasegawa-Johnson, Chang D. YooICLR 2025
Builds on4
- Video as Conditional Graph Hierarchy for Multi-Granular Question AnsweringJunbin Xiao, Angela Yao, Zhiyuan Liu, Yicong Li et al.AAAI 2022 · 145 citations
- Invariant Grounding for Video Question AnsweringYicong Li, Xiang Wang, Junbin Xiao, Wei Ji et al.CVPR 2022 · 108 citations
- Dynamic Graph Representation Learning for Video Dialog via Multi-Modal Shuffled TransformersShijie Geng, Peng Gao, Moitreya Chatterjee, Chiori Hori et al.AAAI 2021 · 50 citations
- Structured Co-reference Graph Attention for Video-grounded DialogueJunyeong Kim, Sunjae Yoon, Dahyun Kim, Chang D. YooAAAI 2021 · 31 citations
Related papers
- Visual Description Grounding Reduces Hallucinations and Boosts Reasoning in LVLMsSreyan Ghosh, Chandra Kiran Reddy Evuru, Sonal Kumar, Utkarsh Tyagi et al.ICLR 2025
- BiST: Bi-directional Spatio-Temporal Reasoning for Video-Grounded DialoguesHung Le, Doyen Sahoo, Nancy F. Chen, Steven C. H. HoiEMNLP 2020 · 30 citations
- VSTAR: A Video-grounded Dialogue Dataset for Situated Semantic Understanding with Scene and Topic TransitionsYuxuan Wang, Zilong Zheng, Xueliang Zhao, Jinpeng Li et al.ACL 2023 · 4 citations
- Let Me Finish My Sentence: Video Temporal Grounding with Holistic Text UnderstandingJongbhin Woo, Hyeonggon Ryu, Youngjoon Jang, Jae-Won Cho et al.ACM MM 2024 · 3 citations
- Learning to Decode Against Compositional Hallucination in Video Multimodal Large Language ModelsWenbin Xing, Quanxing Zha, Lizheng Zu, Mengran Li et al.ICML 2026 · 1 citation
