Violin: A Large-Scale Dataset for Video-and-Language Inference
Jingzhou Liu, Wenhu Chen, Yu Cheng, Zhe Gan, Licheng Yu, Yiming Yang, Jingjing Liu
Abstract
We introduce a new task, Video-and-Language Inference, for joint multimodal understanding of video and text. Given a video clip with aligned subtitles as premise, paired with a natural language hypothesis based on the video content, a model needs to infer whether the hypothesis is entailed or contradicted by the given video clip. A new large-scale dataset, named VIOLIN (VIdeO-and-Language INference), is introduced for this task, which consists of 95,322 videohypothesis pairs from 15,887 video clips, spanning over 582 hours of video. These video clips contain rich content with diverse temporal dynamics, event shifts, and people interactions, collected from two sources: (i) popular TV shows, and (ii) movie clips from YouTube channels. In order to address our new multimodal inference task, a model is required to possess sophisticated reasoning skills, from surface-level grounding (e.g., identifying objects and characters in the video) to in-depth commonsense reasoning (e.g., inferring causal relations of events in the video). We present a detailed analysis of the dataset and an extensive evaluation over many strong baselines, providing valuable insights on the challenges of this new task.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 884dd9d4-508b-4439-8032-cafbee7c3f6cCited by top-tier papers21
- HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-trainingLinjie Li, Yen-Chun Chen, Yu Cheng, Zhe Gan et al.EMNLP 2020 · 387 citations
- Fine-grained Iterative Attention Network for Temporal Language Localization in VideosXiaoye Qu, Pengwei Tang, Zhikang Zou, Yu Cheng et al.ACM MM 2020 · 92 citations
- Long-Form Video-Language Pre-Training with Multimodal Temporal Contrastive LearningYuchong Sun, Hongwei Xue, Ruihua Song, Bei Liu et al.NeurIPS 2022 · 91 citations
- i-Code: An Integrative and Composable Multimodal Learning FrameworkZiyi Yang, Yuwei Fang, Chenguang Zhu, Reid Pryzant et al.AAAI 2023 · 53 citations
- What is More Likely to Happen Next? Video-and-Language Future Event PredictionJie Lei, Licheng Yu, Tamara L. Berg, Mohit BansalEMNLP 2020 · 45 citations
Builds on5
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li et al.ICLR 2020 · 1,825 citations
- VaTeX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language ResearchXin Wang, Jiawei Wu, Jun-Kun Chen, Lei Li et al.ICCV 2019 · 688 citations
- CLEVRER: Collision Events for Video Representation and ReasoningKexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli et al.ICLR 2020 · 584 citations
- Relation-Aware Graph Attention Network for Visual Question AnsweringLinjie Li, Zhe Gan, Yu Cheng, Jingjing LiuICCV 2019 · 391 citations
- TVQA+: Spatio-Temporal Grounding for Video Question AnsweringJie Lei, Licheng Yu, Tamara L. Berg, Mohit BansalACL 2020 · 173 citations
Related papers
- Cross-modal Observation Hypothesis InferenceMengze Li, Kairong Han, Jiahe Xu, Yueying Li et al.ACM MM 2024
- Adaptive Hierarchical Graph Reasoning with Semantic Coherence for Video-and-Language InferenceJuncheng Li, Siliang Tang, Linchao Zhu, Haochen Shi et al.ICCV 2021 · 28 citations
- SynopGround: A Large-Scale Dataset for Multi-Paragraph Video Grounding from TV Dramas and SynopsesChaolei Tan, Zihang Lin, Junfu Pu, Zhongang Qi et al.ACM MM 2024 · 2 citations
- ACQUIRED: A Dataset for Answering Counterfactual Questions In Real-Life VideosTe-Lin Wu, Zi-Yi Dou, Qingyuan Hu, Yu Hou et al.EMNLP 2023 · 3 citations
- InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and GenerationYi Wang, Yinan He, Yizhuo Li, Kunchang Li et al.ICLR 2024 · 467 citations
