Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware Reasoning
SARA GHAZANFARI, Francesco Croce, Nicolas Flammarion, Prashanth Krishnamurthy, Farshad Khorrami, Siddharth Garg
Abstract
Reasoning:
-
Frame 7: The text overlay states, "before reattaching the banana skin to the wall using the same tape." This suggests that the student's action was part of an art installation.
-
Frame 12: The text overlay reads, "but he later added: 'damaging a work of modern art could also be art.'" This implies that the student considered the act of eating the banana as a form of art.
Answer: The student's actions were part of an art installation, and he later explained that damaging a work of modern art could also be considered art. This indicates that he considered the process of eating the banana as an artistic act. B. Because he considered the process of eating a banana is art.
Question: Based on the video, which of the following describes the reason why the student ate the banana? A. Because the banana looks tasty. B. Because he considered the process of eating a banana is art. C. Because he didn't think the banana worth $120,000. D. Because he wanted to followed the man who ate a banana in a exhibition in 2019. … … … … … (a) A Chain-of-Frames reasoning trace generated by our CoF-4B model. 20 30 40 50 60 Accuracy (%) V S I -B e n c h V i d e o -M M E M V B e n c h V i d H a l E v e n t H a l l +3.4% +10.3% +4.9% +7.2% +4.5% +2.7% +2.2% -1.4% +3.8% +6.6%
approach is simple, unified, and self-contained, employing a single-stage inference to handle complex video understanding tasks without relying on auxiliary modules for frame selection or caption generation. For this, we first create COF-DATA, a large dataset of diverse questions, answers, and corresponding frame-grounded reasoning traces from both natural and synthetic videos, spanning various topics and tasks. Our models, obtained fine-tuning video LLMs on this chain-of-frames (CoF) data, generate reasoning traces that accurately identify key frames to answer given questions.
In turn, this consistently improves performance across multiple video understanding benchmarks. Surprisingly, we find that synthetic data alone, despite being out-of-distribution with respect to these real-world benchmarks, provides a significant boost in model accuracy. Code available at GitHub.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 502c87cb-1da7-4234-8db9-2fdd1faba977Cited by top-tier papers9
- TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMsJun Zhang, Teng Wang, Yuying Ge, Yixiao Ge et al.CVPR 2026 · 48 citations
- Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMsChun-Hsiao Yeh, Chenyu Wang, Shengbang Tong, Ta Ying Cheng et al.AAAI 2026 · 35 citations
- SpaceVista: All-Scale Visual Spatial Reasoning from mm to kmPeiwen Sun, Shiqiang Lang, Dongming Wu, Ding Yi et al.ICML 2026 · 19 citations
- VideoAuto-R1: Video Auto Reasoning via Thinking Once, Answering TwiceShuming Liu, Mingchen Zhuge, Changsheng Zhao, Jun Chen et al.CVPR 2026 · 18 citations
- When Can Model-Free Reinforcement Learning be Enough for Thinking?Josiah Hanna, Nicholas CorradoNeurIPS 2025 · 6 citations
Builds on19
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- CLEVRER: Collision Events for Video Representation and ReasoningKexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli et al.ICLR 2020 · 584 citations
- Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language ModelsYushi Hu, Weijia Shi, Xingyu Fu, Dan Roth et al.NeurIPS 2024 · 373 citations
Related papers
- Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric VideosChiara Plizzari, Alessio Tonioni, Yongqin Xian, Achin Kulshrestha et al.CVPR 2025
- MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in VideosKejian Zhu, Zhuoran Jin, Hongbang Yuan, Jiachun Li et al.ICLR 2026 · 22 citations
- Conan: Progressive Learning to Reason Like a Detective over Multi-Scale Visual EvidenceKun Ouyang, Yuanxin Liu, Linli Yao, Yishuo Cai et al.CVPR 2026 · 17 citations
- When Thinking Drifts: Evidential Grounding for Robust Video ReasoningRomy Luo, Zihui Xue, Alex Dimakis, Kristen GraumanNeurIPS 2025 · 21 citations
- CG-Bench: Clue-grounded Question Answering Benchmark for Long Video UnderstandingGuo Chen, Yicheng Liu, Yifei Huang, Baoqi Pei et al.ICLR 2025
