A Graph-Based Framework to Bridge Movies and Synopses
Yu Xiong, Qingqiu Huang, Lingfeng Guo, Hang Zhou, Bolei Zhou, Dahua Lin
Abstract
Inspired by the remarkable advances in video analytics, research teams are stepping towards a greater ambition – movie understanding. However, compared to those activity videos in conventional datasets, movies are significantly different. Generally, movies are much longer and consist of much richer temporal structures. More importantly, the interactions among characters play a central role in expressing the underlying story. To facilitate the efforts along this direction, we construct a dataset called Movie Synopses Associations (MSA) over 327 movies, which provides a synopsis for each movie, together with annotated associations between synopsis paragraphs and movie segments. On top of this dataset, we develop a framework to perform matching between movie segments and synopsis paragraphs. This framework integrates different aspects of a movie, including event dynamics and character interactions, and allows them to be matched with parsed paragraphs, based on a graph-based formulation. Our study shows that the proposed framework remarkably improves the matching accuracy over conventional feature-based methods. It also reveals the importance of narrative structures and character interactions in movie understanding. Dataset and code are available at: https://ycxioooong.github.io/projects/moviesyn
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c7eafbd9-5d15-404f-8bab-437284a3dbffCited by top-tier papers17
- HiT: Hierarchical Transformer with Momentum Contrast for Video-Text RetrievalSong Liu, Haoqi Fan, Shengsheng Qian, Yiru Chen et al.ICCV 2021 · 172 citations
- AutoAD II: The Sequel - Who, When, and What in Movie Audio DescriptionTengda Han, Max Bain, Arsha Nagrani, Gül Varol et al.ICCV 2023 · 55 citations
- Movie Summarization via Sparse Graph ConstructionPinelopi Papalampidi, Frank Keller, Mirella LapataAAAI 2021 · 36 citations
- Learning to Cut by Watching MoviesAlejandro Pardo, Fabian Caba Heilbron, Juan León Alcázar, Ali K. Thabet et al.ICCV 2021 · 26 citations
- Leveraging Weighted Cross-Graph Attention for Visual and Semantic Enhanced Video Captioning NetworkDeepali Verma, Arya Haldar, Tanima DuttaAAAI 2023 · 13 citations
Related papers
- SynopGround: A Large-Scale Dataset for Multi-Paragraph Video Grounding from TV Dramas and SynopsesChaolei Tan, Zihang Lin, Junfu Pu, Zhongang Qi et al.ACM MM 2024 · 2 citations
- VidEvent: A Large Dataset for Understanding Dynamic Evolution of Events in VideosBaoyu Liang, Qile Su, Shoutai Zhu, Yuchen Liang et al.AAAI 2025 · 5 citations
- A Local-to-Global Approach to Multi-Modal Movie Scene SegmentationAnyi Rao, Linning Xu, Yu Xiong, Guodong Xu et al.CVPR 2020
- Multi-view Story Characterization from Movie Plot Synopses and ReviewsSudipta Kar, Gustavo Aguilar, Mirella Lapata, Thamar SolorioEMNLP 2020
- TeViS: Translating Text Synopses to Video StoryboardsXu Gu, Yuchong Sun, Feiyue Ni, Shizhe Chen et al.ACM MM 2023 · 5 citations
