Fine-Grained Captioning of Long Videos through Scene Graph Consolidation
Sanghyeok Chu, Seonguk Seo, Bohyung Han
摘要
Recent advances in vision-language models have led to impressive progress in caption generation for images and short video clips. However, these models remain constrained by their limited temporal receptive fields, making it difficult to produce coherent and comprehensive captions for long videos. While several methods have been proposed to aggregate information across video segments, they often rely on supervised fine-tuning or incur significant computational overhead. To address these challenges, we introduce a novel framework for long video captioning based on graph consolidation. Our approach first generates segment-level captions, corresponding to individual frames or short video intervals, using off-the-shelf visual captioning models. These captions are then parsed into individual scene graphs, which are subsequently consolidated into a unified graph representation that preserves both holistic context and fine-grained details throughout the video. A lightweight graph-to-text decoder then produces the final video-level caption. This framework effectively extends the temporal understanding capabilities of existing models without requiring any additional fine-tuning on long video datasets. Experimental results show that our method significantly outperforms existing LLMbased consolidation approaches, achieving strong zero-shot performance while substantially reducing computational costs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video UnderstandingKe Ma, Jiaqi Tang, Bin Guo, Xueting Han 等ACL 2026
- SegPVSG: Panoptic Video Scene Graph Generation via Temporal Focusing and Generative AugmentationYiKai Li, Quhui Ke, Jinglin Liang, Zhiyuan Zhang 等ICML 2026
- Neuro-Fuzzy Concept Learning for Interpretable Large Multimodal ModelsRitik Mishra, Vanshika Gupta, M. Sajid, M. TanveerICML 2026
它引用的顶会 Paper16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language ModelsMuhammad Maaz, Hanoona Abdul Rasheed, Salman Khan, Fahad KhanACL 2024 · 被引用 279 次
- Streaming Long Video Understanding with Large Language ModelsRui Qian, Xiaoyi Dong, Pan Zhang, Yuhang Zang 等NeurIPS 2024 · 被引用 216 次
相关 Paper
- CALVIN: Improved Contextual Video Captioning via Instruction TuningGowthami Somepalli, Arkabandhu Chowdhury, Jonas Geiping, Ronen Basri 等NeurIPS 2024 · 被引用 4 次
- Discriminative Latent Semantic Graph for Video CaptioningYang Bai, Junyan Wang, Yang Long, Bingzhang Hu 等ACM MM 2021 · 被引用 26 次
- Dense Video Object Captioning from Disjoint SupervisionXingyi Zhou, Anurag Arnab, Chen Sun, Cordelia SchmidICLR 2025
- Scene-VLM: Multimodal Video Scene Segmentation via Vision-Language ModelsNimrod Berman, Adam Botach, Emanuel Ben-Baruch, Shunit Haviv Hakimi 等CVPR 2026 · 被引用 2 次
- Keyframe-Oriented Vision Token Pruning: Enhancing Efficiency of Large Vision Language Models on Long-form Video ProcessingYudong Liu, Jingwei Sun, Yueqian Lin, Jianyi Zhang 等ICCV 2025 · 被引用 22 次
