Unstitching the Chimera: Frame-Level Risk and Train-Free Mitigation for Video Hallucination
Songyuan Yang, Guijian Tang, Kun Hu, Haotian Wang, Shixuan Liu, Wenjing Yang, Long Lan, Huibin Tan
摘要
Hallucination limits the reliability of multimodal large language models (MLLMs), and it is particularly damaging in video where errors manifest as distorted narratives rather than single-frame mistakes. We introduce a frame-first study of Chimera Hallucination: model stitches visual segments that exist in space and time but do not belong to the same event chain, producing a spurious continuous story. We introduce CH-Risk, a single-forward, reference-free risk estimate tailored to this failure mode. CH-Risk combines two complementary signals: SegCoverage@\alpha (\mathrm{SCR}@\alpha\) measures how many event segments are needed to cover most text-to-frame support, exposing long-range stitching; Alignment with Early Temporal Pathway (AETP) measures rank consistency between support and the temporal pathway formed in early–middle layers, exposing stage mismatch. To turn risk into correction, we further propose CH-M(itigation), a train-free two-stage intervention. Segment-aligned Stage-Aligned Frame Routing (sSAFR) re-weights frames before the mid-layer softmax to route attention toward a small set of pathway-aligned segments. Residual Token Calibration (RTC) then stabilizes token usage within selected segments. Extensive experiments across 9 benchmarks and 6 VideoLLMs show that CH-Risk can predict Chimera and that CH-M consistently reduce hallucination and improves task accuracy with negligible overhead (sub-5% latency, sub-2.5% memory, $$1% FLOPs).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper21
- Video-LLaVA: Learning United Visual Representation by Alignment Before ProjectionBin Lin, Yang Ye, Bin Zhu, Jiaxi Cui 等EMNLP 2024 · 被引用 231 次
- Evaluating and Analyzing Relationship Hallucinations in Large Vision-Language ModelsMingrui Wu, Jiayi Ji, Oucheng Huang, Jiale Li 等ICML 2024 · 被引用 32 次
- LVBench: An Extreme Long Video Understanding BenchmarkWeihan Wang, Zehai He, Wenyi Hong, Yean Cheng 等ICCV 2025 · 被引用 28 次
- Vista-llama: Reducing Hallucination in Video Language Models via Equal Distance to Visual TokensFan Ma, Xiaojie Jin, Heng Wang, Yuchen Xian 等CVPR 2024 · 被引用 11 次
- Map the Flow: Revealing Hidden Pathways of Information in VideoLLMsMinji Kim, Taekyung Kim, Bohyung HanICLR 2026 · 被引用 8 次
相关 Paper
- SmartSight: Mitigating Hallucination in Video-LLMs Without Compromising Video Understanding via Temporal Attention CollapseYiming Sun, Mi Zhang, Feifei Li, Geng Hong 等AAAI 2026 · 被引用 5 次
- One Token, Two Fates: A Unified Framework via Vision Token Manipulation Against MLLMs HallucinationZhan Fa, Yue Duan, Jian Zhang, Lei Qi 等CVPR 2026 · 被引用 2 次
- Prefill-Time Intervention for Mitigating Hallucination in Large Vision-Language ModelsChengsheng Zhang, Chenghao Sun, Xinyan Jiang, Wei Li 等CVPR 2026 · 被引用 2 次
- Grounding or Guessing? Visual Signals for Detecting Hallucinations in Sign Language TranslationYasser Hamidullah, Koel Dutta Chowdhury, Yusser Al Ghussin, Shakib Yazdani 等ICLR 2026 · 被引用 3 次
- Vision-Language Introspection: Mitigating Overconfident Hallucinations in MLLMs via Interpretable Bi-Causal SteeringShuliang Liu, Songbo Yang, Dong Fang, Sihang Jia 等ACL 2026 · 被引用 9 次
