Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware Reasoning
SARA GHAZANFARI, Francesco Croce, Nicolas Flammarion, Prashanth Krishnamurthy, Farshad Khorrami, Siddharth Garg
摘要
Reasoning:
-
Frame 7: The text overlay states, "before reattaching the banana skin to the wall using the same tape." This suggests that the student's action was part of an art installation.
-
Frame 12: The text overlay reads, "but he later added: 'damaging a work of modern art could also be art.'" This implies that the student considered the act of eating the banana as a form of art.
Answer: The student's actions were part of an art installation, and he later explained that damaging a work of modern art could also be considered art. This indicates that he considered the process of eating the banana as an artistic act. B. Because he considered the process of eating a banana is art.
Question: Based on the video, which of the following describes the reason why the student ate the banana? A. Because the banana looks tasty. B. Because he considered the process of eating a banana is art. C. Because he didn't think the banana worth $120,000. D. Because he wanted to followed the man who ate a banana in a exhibition in 2019. … … … … … (a) A Chain-of-Frames reasoning trace generated by our CoF-4B model. 20 30 40 50 60 Accuracy (%) V S I -B e n c h V i d e o -M M E M V B e n c h V i d H a l E v e n t H a l l +3.4% +10.3% +4.9% +7.2% +4.5% +2.7% +2.2% -1.4% +3.8% +6.6%
approach is simple, unified, and self-contained, employing a single-stage inference to handle complex video understanding tasks without relying on auxiliary modules for frame selection or caption generation. For this, we first create COF-DATA, a large dataset of diverse questions, answers, and corresponding frame-grounded reasoning traces from both natural and synthetic videos, spanning various topics and tasks. Our models, obtained fine-tuning video LLMs on this chain-of-frames (CoF) data, generate reasoning traces that accurately identify key frames to answer given questions.
In turn, this consistently improves performance across multiple video understanding benchmarks. Surprisingly, we find that synthetic data alone, despite being out-of-distribution with respect to these real-world benchmarks, provides a significant boost in model accuracy. Code available at GitHub.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMsJun Zhang, Teng Wang, Yuying Ge, Yixiao Ge 等CVPR 2026 · 被引用 48 次
- Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMsChun-Hsiao Yeh, Chenyu Wang, Shengbang Tong, Ta Ying Cheng 等AAAI 2026 · 被引用 35 次
- SpaceVista: All-Scale Visual Spatial Reasoning from mm to kmPeiwen Sun, Shiqiang Lang, Dongming Wu, Ding Yi 等ICML 2026 · 被引用 19 次
- VideoAuto-R1: Video Auto Reasoning via Thinking Once, Answering TwiceShuming Liu, Mingchen Zhuge, Changsheng Zhao, Jun Chen 等CVPR 2026 · 被引用 18 次
- When Can Model-Free Reinforcement Learning be Enough for Thinking?Josiah Hanna, Nicholas CorradoNeurIPS 2025 · 被引用 6 次
它引用的顶会 Paper19
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- CLEVRER: Collision Events for Video Representation and ReasoningKexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli 等ICLR 2020 · 被引用 584 次
- Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language ModelsYushi Hu, Weijia Shi, Xingyu Fu, Dan Roth 等NeurIPS 2024 · 被引用 373 次
相关 Paper
- Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric VideosChiara Plizzari, Alessio Tonioni, Yongqin Xian, Achin Kulshrestha 等CVPR 2025
- MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in VideosKejian Zhu, Zhuoran Jin, Hongbang Yuan, Jiachun Li 等ICLR 2026 · 被引用 22 次
- Conan: Progressive Learning to Reason Like a Detective over Multi-Scale Visual EvidenceKun Ouyang, Yuanxin Liu, Linli Yao, Yishuo Cai 等CVPR 2026 · 被引用 17 次
- When Thinking Drifts: Evidential Grounding for Robust Video ReasoningRomy Luo, Zihui Xue, Alex Dimakis, Kristen GraumanNeurIPS 2025 · 被引用 21 次
- CG-Bench: Clue-grounded Question Answering Benchmark for Long Video UnderstandingGuo Chen, Yicheng Liu, Yifei Huang, Baoqi Pei 等ICLR 2025
