MASH-VLM: Mitigating Action-Scene Hallucination in Video-LLMs through Disentangled Spatial-Temporal Representations
Kyungho Bae, Jinhyung Kim, Sihaeng Lee, Soonyoung Lee, Gunhee Lee, Jinwoo Choi
Abstract
Yes, the background in the video is a boxing ring
No, the background in the video is not a boxing ring. It is a room with books and shelves.
Question: Is the background in the video a boxing ring?
The video shows a person skiing on a mountain. The person is wearing skis and is riding on a slope.
The video shows a snow-covered mountain slope with no one in sight. Sunlight casts a warm glow over the serene landscape, and immersive experience.
Yes, two men are throwing punches in a boxing ring
The video starts with a white title screen. We then see a skier who is snowboarding on a steep mountain slope covered in snow
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 16e0d5a1-6775-4dfe-b48f-c4f3266bfc3cCited by top-tier papers8
- Vision-Language Introspection: Mitigating Overconfident Hallucinations in MLLMs via Interpretable Bi-Causal SteeringShuliang Liu, Songbo Yang, Dong Fang, Sihang Jia et al.ACL 2026 · 9 citations
- SEASON: Mitigating Temporal Hallucination in Video Large Language Models via Self-Diagnostic Contrastive DecodingChang-Hsun Wu, Kai-Po Chang, Yu-Yang Sheng, Hung-Kai Chung et al.CVPR 2026 · 8 citations
- Mitigating Hallucination in VideoLLMs via Temporal-Aware Activation EngineeringJianfeng Cai, Jiale Hong, Zongmeng Zhang, Wengang Zhou et al.NeurIPS 2025 · 7 citations
- INFACT: A Diagnostic Benchmark for Induced Faithfulness and Factuality Hallucinations in Video-LLMsJunqi Yang, Yuecong Min, Jie Zhang, Shiguang Shan et al.ACL 2026 · 7 citations
- SmartSight: Mitigating Hallucination in Video-LLMs Without Compromising Video Understanding via Temporal Attention CollapseYiming Sun, Mi Zhang, Feifei Li, Geng Hong et al.AAAI 2026 · 5 citations
Builds on23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
Related papers
- Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video ContentQiuheng Wang, Yukai Shi, Jiarong Ou, Rui Chen et al.CVPR 2025
- PreciseCam: Precise Camera Control for Text-to-Image GenerationEdurne Bernal-Berdun, Ana Serrano, Belén Masiá, Matheus Gadelha et al.CVPR 2025
- DualVector: Unsupervised Vector Font Synthesis with Dual-Part RepresentationYing-Tian Liu, Zhifei Zhang, Yuan-Chen Guo, Matthew Fisher et al.CVPR 2023
- UFVideo: Towards Unified Fine-Grained Video Cooperative Understanding with Large Language ModelsHewen Pan, Cong Wei, Dashuang Liang, Zepeng Huang et al.CVPR 2026 · 3 citations
- Video Spatial Reasoning with Object-Centric 3D RolloutHaoran Tang, Meng Cao, Ruyang Liu, Xiaoxi Liang et al.AAAI 2026 · 3 citations
