MEGA: Multimodal Alignment Aggregation and Distillation For Cinematic Video Segmentation
Najmeh Sadoughi, Xinyu Li, Avijit Vajpayee, David Fan, Bing Shuai, Hector J. Santos-Villalobos, Vimal Bhat, Rohith MV
Abstract
Previous research has studied the task of segmenting cinematic videos into scenes and into narrative acts. However, these studies have overlooked the essential task of multimodal alignment and fusion for effectively and efficiently processing long-form videos (> 60min). In this paper, we introduce Multimodal alignmEnt aGgregation and distillAtion (MEGA) for cinematic long-video segmentation. MEGA tackles the challenge by leveraging multiple media modalities. The method coarsely aligns inputs of variable lengths and different modalities with alignment positional encoding. To maintain temporal synchronization while reducing computation, we further introduce an enhanced bottleneck fusion layer which uses temporal alignment. Additionally, MEGA employs a novel contrastive loss to synchronize and transfer labels across modalities, enabling act segmentation from labeled synopsis sentences on video shots. Our experimental results show that MEGA outperforms state-of-the-art methods on MovieNet dataset for scene segmentation (with an Average Precision improvement of +1.19%) and on TRI-POD dataset for act segmentation (with a Total Agreement improvement of +5.51%).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b82dd901-e31d-4c53-91e3-2419f4f0f57dCited by top-tier papers2
- Chapter-Llama: Efficient Chaptering in Hour-Long Videos with LLMsLucas Ventura, Antoine Yang, Cordelia Schmid, Gül VarolCVPR 2025
- Open Your Model's Eyes: Video and Context-Aware Multimodal Backchannel PredictionMin-Jae Kim, Jun-Yeong Moon, Mujeen Sung, Gyeong-Moon ParkACL 2026
Builds on15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 2,258 citations
- Perceiver: General Perception with Iterative AttentionAndrew Jaegle, Felix Gimeno, Andy Brock, Oriol Vinyals et al.ICML 2021 · 1,399 citations
- Attention Bottlenecks for Multimodal FusionArsha Nagrani, Shan Yang, Anurag Arnab, Aren Jansen et al.NeurIPS 2021 · 884 citations
Related papers
- Long-Form Video-Language Pre-Training with Multimodal Temporal Contrastive LearningYuchong Sun, Hongwei Xue, Ruihua Song, Bei Liu et al.NeurIPS 2022 · 91 citations
- Modeling Motion with Multi-Modal Features for Text-Based Video SegmentationWangbo Zhao, Kai Wang, Xiangxiang Chu, Fuzhao Xue et al.CVPR 2022 · 23 citations
- SynopGround: A Large-Scale Dataset for Multi-Paragraph Video Grounding from TV Dramas and SynopsesChaolei Tan, Zihang Lin, Junfu Pu, Zhongang Qi et al.ACM MM 2024 · 2 citations
- CONE: An Efficient COarse-to-fiNE Alignment Framework for Long Video Temporal GroundingZhijian Hou, Wanjun Zhong, Lei Ji, Difei Gao et al.ACL 2023 · 17 citations
- A Local-to-Global Approach to Multi-Modal Movie Scene SegmentationAnyi Rao, Linning Xu, Yu Xiong, Guodong Xu et al.CVPR 2020
