ICML2026
Chain-of-Glimpse: Search-Guided Progressive Object-Grounded Reasoning for Video Understanding
ZHIXUAN WU, Quanxing Zha, Teng Wang, Genbao Xu, Wenyuan Gu, Wei Rao, Nan Ma, Bo Cheng, Soujanya Poria
被引用 1 次
摘要
Video understanding requires identifying and reasoning over semantically discriminative visual objects across frames, yet existing object-agnostic solutions struggle to effectively handle substantial object variations over time. To address this, we introduce Chain-of-Glimpse, a search-guided progressive object-grounded reasoning framework that explicitly anchors each reasoning step to specific visual evidence regions, enabling compositional and multi-step decision-making. Formally, Chain-of-Glimpse formulates video reasoning as a step-by-step process that incrementally builds spatially grounded traces around task-relevant visual objects, thereby mitigating over-reliance on saliency-driven cues. Specifically, Chain-of-Glimpse features a search-guided controller, optimized via reinforcement learning with a format reward that significantly incentivizes grounding capability, to iteratively ground visual evidence regions and form reliable reasoning trajectories, yielding accurate and interpretable multi-step decisions. Extensive evaluations across two categories of video reasoning benchmarks, including general video reasoning benchmarks such as NEx-TQA, Video-Holmes, CG-Bench-Reasoning, and VRBench, and grounded video reasoning benchmarks such as NExT-GQA, demonstrate that Chain-of-Glimpse consistently improves performance while exhibiting strong robustness and generalization across diverse video reasoning tasks.