Lune

ICML2026Top-tier venue

Video-in-the-Loop: Span-Grounded Long Video QA with Interleaved Reasoning

Chendong Wang, Donglin Bai, Yifan Yang, Xiao Jin, Anlan Zhang, Rui Wang, Shiqi Jiang, Yuqing Yang, Hao Wu, Qi Dai, Chong Luo, Ting Cao

2026Year
3Citations
1Top-tier citations

Abstract

We present Video-in-the-Loop\textit{Video-in-the-Loop} (ViTL), a two-stage long-video QA framework that preserves a fixed token budget by first localizing\textit{localizing} question-relevant interval(s) with a low-fps skim and then answering\textit{answering} via span-aware reallocation of visual tokens at higher effective frame rate, emitting an interleaved output with both spans and the final option for direct attribution. We also introduce VGrounding-QA\textit{VGrounding-QA}, which converts description based event graphs into span-grounded\textit{span-grounded} multiple-choice QA by pairing each question with ground-truth\textit{ground-truth} time span(s) and related reasoning. ViTL is trained end-to-end with an interleaved group-relative objective that couples temporal IoU for localization with answer correctness, allowing credit to flow from answers back to spans without increasing compute. Under fixed token budgets, ViTL attains up to 8.6% with 50% less frame input on long-video QA and temporal grounding (e.g., Charades-STA, ActivityNet-Captions) and ablations show that span-aware token reallocation consistently surpasses uniform sampling. Together, VGrounding-QA\textit{VGrounding-QA} and ViTL provide an interpretable, compute-efficient recipe for scalable long-video QA.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

Cited by top-tier papers1

Ask how each one uses it

Builds on14

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines