Spoken Moments: Learning Joint Audio-Visual Representations From Video Descriptions
Mathew Monfort, SouYoung Jin, Alexander H. Liu, David Harwath, Rogério Feris, James R. Glass, Aude Oliva
Abstract
When people observe events, they are able to abstract key information and build concise summaries of what is happening. These summaries include contextual and semantic information describing the important high-level details (what, where, who and how) of the observed event and exclude background information that is deemed unimportant to the observer. With this in mind, the descriptions people generate for videos of different dynamic events can greatly improve our understanding of the key information of interest in each video. These descriptions can be captured in captions that provide expanded attributes for video labeling (e.g. actions/objects/scenes/sentiment/etc.) while allowing us to gain new insight into what people find important or necessary to summarize specific events. Existing caption datasets for video understanding are either small in scale or restricted to a specific domain. To address this, we present the Spoken Moments (S-MiT) dataset of 500k spoken captions each attributed to a unique short video depicting a broad range of different events. We collect our descriptions using audio recordings to ensure that they remain as natural and concise as possible while allowing us to scale the size of a large classification dataset. In order to utilize our proposed dataset, we present a novel Adaptive Mean Margin (AMM) approach to contrastive learning and evaluate our models on video/caption retrieval on multiple datasets. We show that our AMM approach consistently improves our results and that models trained on our Spoken Moments dataset generalize better than those trained on other videocaption datasets. http://moments.csail.mit.edu/spoken.html
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 72bf6492-da19-400a-89e7-96405deb287aCited by top-tier papers31
- Learning Interactive Real-World SimulatorsSherry Yang, Yilun Du, Seyed Kamyar Seyed Ghasemipour, Jonathan Tompson et al.ICLR 2024 · 399 citations
- VideoChat-Flash: Hierarchical Compression for Long-Context Video ModelingXinhao Li, Yi Wang, Jiashuo Yu, Xiangyu Zeng et al.ICLR 2026 · 172 citations
- PerceptionLM: Open-Access Data and Models for Detailed Visual UnderstandingJang Hyun Cho, Andrea Madotto, Effrosyni Mavroudi, Triantafyllos Afouras et al.NeurIPS 2025 · 97 citations
- Verbs in Action: Improving verb understanding in video-language modelsLiliane Momeni, Mathilde Caron, Arsha Nagrani, Andrew Zisserman et al.ICCV 2023 · 93 citations
- VideoPrism: A Foundational Visual Encoder for Video UnderstandingLong Zhao, Nitesh Bharadwaj Gundavarapu, Liangzhe Yuan, Hao Zhou et al.ICML 2024 · 91 citations
Builds on6
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi et al.ICCV 2019 · 1,437 citations
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy et al.ICCV 2019 · 1,396 citations
- VaTeX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language ResearchXin Wang, Jiawei Wu, Jun-Kun Chen, Lei Li et al.ICCV 2019 · 688 citations
Related papers
- Hierarchical Video-Moment Retrieval and Step-CaptioningAbhay Zala, Jaemin Cho, Satwik Kottur, Xilun Chen et al.CVPR 2023
- Auto-ACD: A Large-scale Dataset for Audio-Language Representation LearningLuoyi Sun, Xuenan Xu, Mengyue Wu, Weidi XieACM MM 2024 · 23 citations
- Youku Dense Caption: A Large-scale Chinese Video Dense Caption Dataset and BenchmarksZixuan Xiong, Guangwei Xu, Wenkai Zhang, Yuan Miao et al.ICLR 2025
- Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual AwarenessJiaxing Zhao, Boyuan Sun, Xiang Chen, Xihan WeiAAAI 2026
- HACS: Human Action Clips and Segments Dataset for Recognition and Temporal LocalizationHang Zhao, Antonio Torralba, Lorenzo Torresani, Zhicheng YanICCV 2019 · 298 citations
