A Graph-Based Framework to Bridge Movies and Synopses
Yu Xiong, Qingqiu Huang, Lingfeng Guo, Hang Zhou, Bolei Zhou, Dahua Lin
摘要
Inspired by the remarkable advances in video analytics, research teams are stepping towards a greater ambition – movie understanding. However, compared to those activity videos in conventional datasets, movies are significantly different. Generally, movies are much longer and consist of much richer temporal structures. More importantly, the interactions among characters play a central role in expressing the underlying story. To facilitate the efforts along this direction, we construct a dataset called Movie Synopses Associations (MSA) over 327 movies, which provides a synopsis for each movie, together with annotated associations between synopsis paragraphs and movie segments. On top of this dataset, we develop a framework to perform matching between movie segments and synopsis paragraphs. This framework integrates different aspects of a movie, including event dynamics and character interactions, and allows them to be matched with parsed paragraphs, based on a graph-based formulation. Our study shows that the proposed framework remarkably improves the matching accuracy over conventional feature-based methods. It also reveals the importance of narrative structures and character interactions in movie understanding. Dataset and code are available at: https://ycxioooong.github.io/projects/moviesyn
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- HiT: Hierarchical Transformer with Momentum Contrast for Video-Text RetrievalSong Liu, Haoqi Fan, Shengsheng Qian, Yiru Chen 等ICCV 2021 · 被引用 172 次
- AutoAD II: The Sequel - Who, When, and What in Movie Audio DescriptionTengda Han, Max Bain, Arsha Nagrani, Gül Varol 等ICCV 2023 · 被引用 55 次
- Movie Summarization via Sparse Graph ConstructionPinelopi Papalampidi, Frank Keller, Mirella LapataAAAI 2021 · 被引用 36 次
- Learning to Cut by Watching MoviesAlejandro Pardo, Fabian Caba Heilbron, Juan León Alcázar, Ali K. Thabet 等ICCV 2021 · 被引用 26 次
- Leveraging Weighted Cross-Graph Attention for Visual and Semantic Enhanced Video Captioning NetworkDeepali Verma, Arya Haldar, Tanima DuttaAAAI 2023 · 被引用 13 次
相关 Paper
- SynopGround: A Large-Scale Dataset for Multi-Paragraph Video Grounding from TV Dramas and SynopsesChaolei Tan, Zihang Lin, Junfu Pu, Zhongang Qi 等ACM MM 2024 · 被引用 2 次
- VidEvent: A Large Dataset for Understanding Dynamic Evolution of Events in VideosBaoyu Liang, Qile Su, Shoutai Zhu, Yuchen Liang 等AAAI 2025 · 被引用 5 次
- A Local-to-Global Approach to Multi-Modal Movie Scene SegmentationAnyi Rao, Linning Xu, Yu Xiong, Guodong Xu 等CVPR 2020
- Multi-view Story Characterization from Movie Plot Synopses and ReviewsSudipta Kar, Gustavo Aguilar, Mirella Lapata, Thamar SolorioEMNLP 2020
- TeViS: Translating Text Synopses to Video StoryboardsXu Gu, Yuchong Sun, Feiyue Ni, Shizhe Chen 等ACM MM 2023 · 被引用 5 次
