Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos
Mingfei Han, Linjie Yang, Xiaojun Chang, Lina Yao, Heng Wang
Abstract
A short clip of video may contain progression of multiple events and an interesting story line. A human need to capture both the event in every shot and associate them together to understand the story behind it. In this work, we present a new multi-shot video understanding benchmark Shot2Story with detailed shot-level captions, comprehensive video summaries and question-answering pairs. To facilitate better semantic understanding of videos, we provide captions for both visual signals and human narrations. We design several distinct tasks including single-shot video captioning, multi-shot video summarization, and multi-shot video question answering. Preliminary experiments show some challenges to generate a long and comprehensive video summary for multi-shot videos. Nevertheless, the generated imperfect summaries can already achieve competitive performance on existing video understanding tasks such as video question-answering, promoting an under-explored setting of video understanding with detailed summaries.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 71e75c01-414a-46cc-943b-69f314190d32Cited by top-tier papers14
- Eagle 2.5: Boosting Long-Context Post-Training for Frontier Vision-Language ModelsGuo Chen, Zhiqi Li, Shihao Wang, Jindong Jiang et al.NeurIPS 2025 · 69 citations
- StreamBridge: Turning Your Offline Video Large Language Model into a Proactive Streaming AssistantHaibo Wang, Bo Feng, Zhengfeng Lai, Mingze Xu et al.NeurIPS 2025 · 63 citations
- Streaming Video Instruction TuningJiaer Xia, Peixian Chen, Mengdan Zhang, Xing Sun et al.CVPR 2026 · 28 citations
- LiveStar: Live Streaming Assistant for Real-World Online Video UnderstandingZhenyu Yang, Kairui Zhang, Yuhang Hu, Bing Wang et al.NeurIPS 2025 · 26 citations
- Advancing Complex Video Object Segmentation via Progressive Concept ConstructionZhixiong Zhang, Shuangrui Ding, Xiaoyi Dong, Songxin He et al.ICLR 2026 · 17 citations
Builds on11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis et al.CVPR 2022 · 525 citations
- Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language ModelsMuhammad Maaz, Hanoona Abdul Rasheed, Salman Khan, Fahad KhanACL 2024 · 279 citations
Related papers
- OneStory: Coherent Multi-Shot Video Generation with Adaptive MemoryZhaochong An, Menglin Jia, Haonan Qiu, Zijian Zhou et al.CVPR 2026 · 33 citations
- Synchronized Video Storytelling: Generating Video Narrations with Structured StorylineDingyi Yang, Chunru Zhan, Ziheng Wang, Biao Wang et al.ACL 2024
- MVBench: A Comprehensive Multi-modal Video Understanding BenchmarkKunchang Li, Yali Wang, Yinan He, Yizhuo Li et al.CVPR 2024
- ScaleLong: A Multi-Timescale Benchmark for Long Video UnderstandingDavid Ma, Huaqing Yuan, Xingjian Wang, Qianbo Zang et al.ICLR 2026 · 7 citations
- VidComposition: Can MLLMs Analyze Compositions in Compiled Videos?Yunlong Tang, Junjia Guo, Hang Hua, Susan Liang et al.CVPR 2025
