Lune

CVPR2026Top-tier venue

Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware Reasoning

SARA GHAZANFARI, Francesco Croce, Nicolas Flammarion, Prashanth Krishnamurthy, Farshad Khorrami, Siddharth Garg

2026Year
35Citations
9Top-tier citations

Abstract

Reasoning:

  1. Frame 7: The text overlay states, "before reattaching the banana skin to the wall using the same tape." This suggests that the student's action was part of an art installation.

  2. Frame 12: The text overlay reads, "but he later added: 'damaging a work of modern art could also be art.'" This implies that the student considered the act of eating the banana as a form of art.

Answer: The student's actions were part of an art installation, and he later explained that damaging a work of modern art could also be considered art. This indicates that he considered the process of eating the banana as an artistic act. B. Because he considered the process of eating a banana is art.

Question: Based on the video, which of the following describes the reason why the student ate the banana? A. Because the banana looks tasty. B. Because he considered the process of eating a banana is art. C. Because he didn't think the banana worth $120,000. D. Because he wanted to followed the man who ate a banana in a exhibition in 2019. … … … … … (a) A Chain-of-Frames reasoning trace generated by our CoF-4B model. 20 30 40 50 60 Accuracy (%) V S I -B e n c h V i d e o -M M E M V B e n c h V i d H a l E v e n t H a l l +3.4% +10.3% +4.9% +7.2% +4.5% +2.7% +2.2% -1.4% +3.8% +6.6%

approach is simple, unified, and self-contained, employing a single-stage inference to handle complex video understanding tasks without relying on auxiliary modules for frame selection or caption generation. For this, we first create COF-DATA, a large dataset of diverse questions, answers, and corresponding frame-grounded reasoning traces from both natural and synthetic videos, spanning various topics and tasks. Our models, obtained fine-tuning video LLMs on this chain-of-frames (CoF) data, generate reasoning traces that accurately identify key frames to answer given questions.

In turn, this consistently improves performance across multiple video understanding benchmarks. Surprisingly, we find that synthetic data alone, despite being out-of-distribution with respect to these real-world benchmarks, provides a significant boost in model accuracy. Code available at GitHub.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 502c87cb-1da7-4234-8db9-2fdd1faba977

Cited by top-tier papers9

Ask how each one uses it

Builds on19

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines