3MASSIV: Multilingual, Multimodal and Multi-Aspect dataset of Social Media Short Videos
Vikram Gupta, Trisha Mittal, Puneet Mathur, Vaibhav Mishra, Mayank Maheshwari, Aniket Bera, Debdoot Mukherjee, Dinesh Manocha
Abstract
We present 3MASSIV, a multilingual, multimodal and multi-aspect, expertly-annotated dataset of diverse short videos extracted from short-video social media platform - Moj. 3MASSIV comprises of 50k short videos (20 seconds average duration) and 100K unlabeled videos in 11 different languages and captures popular short video trends like pranks, fails, romance, comedy expressed via unique audio-visual formats like self-shot videos, reaction videos, lip-synching, self-sung songs, etc. 3MASSIV presents an opportunity for multimodal and multilingual semantic understanding on these unique videos by annotating them for concepts, affective states, media types, and audio language. We present a thorough analysis of 3MASSIV and highlight the variety and unique aspects of our dataset compared to other contemporary popular datasets with strong baselines. We also show how the social media content in 3MASSIV is dynamic and temporal in nature, which can be used for semantic understanding tasks and cross-lingual analysis.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7d43170f-e89a-4a91-afbd-460717bda4e4Cited by top-tier papers2
- Joyful: Joint Modality Fusion and Graph Contrastive Learning for Multimoda Emotion RecognitionDongyuan Li, Yusong Wang, Kotaro Funakoshi, Manabu OkumuraEMNLP 2023 · 46 citations
- Video Recognition in Portrait ModeMingfei Han, Linjie Yang, Xiaojie Jin, Jiashi Feng et al.CVPR 2024
Builds on8
- VaTeX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language ResearchXin Wang, Jiawei Wu, Jun-Kun Chen, Lei Li et al.ICCV 2019 · 688 citations
- HACS: Human Action Clips and Segments Dataset for Recognition and Temporal LocalizationHang Zhao, Antonio Torralba, Lorenzo Torresani, Zhicheng YanICCV 2019 · 298 citations
- Parameter-Efficient Transfer from Sequential Behaviors for User Modeling and RecommendationFajie Yuan, Xiangnan He, Alexandros Karatzoglou, Liguang ZhangSIGIR 2020 · 155 citations
- MMAct: A Large-Scale Dataset for Cross Modal Human Action UnderstandingQuan Kong, Ziming Wu, Ziwei Deng, Martin Klinkigt et al.ICCV 2019 · 108 citations
- SVD: A Large-Scale Short Video Dataset for Near-Duplicate Video RetrievalQing-Yuan Jiang, Yi He, Gen Li, Jian Lin et al.ICCV 2019 · 52 citations
Related papers
- A Large-Scale Dataset for Short-Video Topic Peak Prediction and a Large Heterogeneous Graph ModelShangheng Chen, Shengsheng Qian, Quan Fang, Jun Hu et al.ACM MM 2025
- Spatiotemporal Fine-grained Video Description for Short VideosTe Yang, Jian Jia, Bo Wang, Yanhua Cheng et al.ACM MM 2024 · 1 citation
- MASIVE: Open-Ended Affective State Identification in English and SpanishNicholas Deas, Elsbeth Turcan, Iván Pérez Mejía, Kathleen R. McKeownEMNLP 2024 · 3 citations
- CMU-MOSEAS: A Multimodal Language Dataset for Spanish, Portuguese, German and FrenchAmirAli Bagher Zadeh, Yansheng Cao, Smon Hessner, Paul Pu Liang et al.EMNLP 2020 · 43 citations
- M³AV: A Multimodal, Multigenre, and Multipurpose Audio-Visual Academic Lecture DatasetZhe Chen, Heyang Liu, Wenyi Yu, Guangzhi Sun et al.ACL 2024 · 2 citations
