MAD: A Scalable Dataset for Language Grounding in Videos from Movie Audio Descriptions
Mattia Soldan, Alejandro Pardo, Juan León Alcázar, Fabian Caba Heilbron, Chen Zhao, Silvio Giancola, Bernard Ghanem
Abstract
The recent and increasing interest in video-language research has driven the development of large-scale datasets that enable data-intensive machine learning techniques. In comparison, limited effort has been made at assessing the fitness of these datasets for the video-language grounding task. Recent works have begun to discover significant limitations in these datasets, suggesting that state-of-the-art techniques commonly overfit to hidden dataset biases. In this work, we present MAD (Movie Audio Descriptions), a novel benchmark that departs from the paradigm of augmenting existing video datasets with text annotations and focuses on crawling and aligning available audio descriptions of mainstream movies. MAD contains over 384, 000 natural language sentences grounded in over 1, 200 hours of videos and exhibits a significant reduction in the currently diagnosed biases for video-language grounding datasets. MAD's collection strategy enables a novel and more challenging version of video-language grounding, where short temporal moments (typically seconds long) must be accurately grounded in diverse long-form videos that can last up to three hours. We have released MAD's data and baselines code at https://github.com/Soldelli/MAD.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers57
- TMR: Text-to-Motion Retrieval Using Contrastive 3D Human Motion SynthesisMathis Petrovich, Michael J. Black, Gül VarolICCV 2023 · 192 citations
- MomentDiff: Generative Video Moment Retrieval from Random to RealPandeng Li, Chen-Wei Xie, Hongtao Xie, Liming Zhao et al.NeurIPS 2023 · 113 citations
- MovieChat: From Dense Token to Sparse Memory for Long Video UnderstandingEnxin Song, Wenhao Chai, Guanhong Wang, Yucheng Zhang et al.CVPR 2024 · 95 citations
- Connecting Multi-modal Contrastive RepresentationsZehan Wang, Yang Zhao, Xize Cheng, Haifeng Huang et al.NeurIPS 2023 · 60 citations
- AutoAD II: The Sequel - Who, When, and What in Movie Audio DescriptionTengda Han, Max Bain, Arsha Nagrani, Gül Varol et al.ICCV 2023 · 55 citations
Builds on13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- DeepGCNs: Can GCNs Go As Deep As CNNs?Guohao Li, Matthias Müller, Ali K. Thabet, Bernard GhanemICCV 2019 · 1,586 citations
- Learning 2D Temporal Adjacent Networks for Moment Localization with Natural LanguageSongyang Zhang, Houwen Peng, Jianlong Fu, Jiebo LuoAAAI 2020 · 579 citations
- Video Self-Stitching Graph Network for Temporal Action LocalizationChen Zhao, Ali K. Thabet, Bernard GhanemICCV 2021 · 179 citations
- Jointly Cross- and Self-Modal Graph Attention Network for Query-Based Moment LocalizationDaizong Liu, Xiaoye Qu, Xiao-Yang Liu, Jianfeng Dong et al.ACM MM 2020 · 115 citations
Related papers
- Localizing Moments in Long Video Via Multimodal GuidanceWayner Barrios, Mattia Soldan, Alberto Mario Ceballos-Arroyo, Fabian Caba Heilbron et al.ICCV 2023 · 32 citations
- SnAG: Scalable and Accurate Video GroundingFangzhou Mu, Sicheng Mo, Yin LiCVPR 2024 · 13 citations
- VidLA: Video-Language Alignment at ScaleMamshad Nayeem Rizve, Fan Fei, Jayakrishnan Unnikrishnan, Son Tran et al.CVPR 2024 · 3 citations
- VALU: A Benchmark for Video Anomaly Temporal Localization and Understanding at Multiple Semantic LevelsYixiao He, Menghao Zhang, Haifeng Sun, Jing Wang et al.ACL 2026
- Orchestrating Audio: Multi-Agent Framework for Long-Video Audio SynthesisYehang Zhang, Xinli Xu, Xiaojie Xu, Doudou Zhang et al.EMNLP 2025
