Violin: A Large-Scale Dataset for Video-and-Language Inference
Jingzhou Liu, Wenhu Chen, Yu Cheng, Zhe Gan, Licheng Yu, Yiming Yang, Jingjing Liu
摘要
We introduce a new task, Video-and-Language Inference, for joint multimodal understanding of video and text. Given a video clip with aligned subtitles as premise, paired with a natural language hypothesis based on the video content, a model needs to infer whether the hypothesis is entailed or contradicted by the given video clip. A new large-scale dataset, named VIOLIN (VIdeO-and-Language INference), is introduced for this task, which consists of 95,322 videohypothesis pairs from 15,887 video clips, spanning over 582 hours of video. These video clips contain rich content with diverse temporal dynamics, event shifts, and people interactions, collected from two sources: (i) popular TV shows, and (ii) movie clips from YouTube channels. In order to address our new multimodal inference task, a model is required to possess sophisticated reasoning skills, from surface-level grounding (e.g., identifying objects and characters in the video) to in-depth commonsense reasoning (e.g., inferring causal relations of events in the video). We present a detailed analysis of the dataset and an extensive evaluation over many strong baselines, providing valuable insights on the challenges of this new task.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-trainingLinjie Li, Yen-Chun Chen, Yu Cheng, Zhe Gan 等EMNLP 2020 · 被引用 387 次
- Fine-grained Iterative Attention Network for Temporal Language Localization in VideosXiaoye Qu, Pengwei Tang, Zhikang Zou, Yu Cheng 等ACM MM 2020 · 被引用 92 次
- Long-Form Video-Language Pre-Training with Multimodal Temporal Contrastive LearningYuchong Sun, Hongwei Xue, Ruihua Song, Bei Liu 等NeurIPS 2022 · 被引用 91 次
- i-Code: An Integrative and Composable Multimodal Learning FrameworkZiyi Yang, Yuwei Fang, Chenguang Zhu, Reid Pryzant 等AAAI 2023 · 被引用 53 次
- What is More Likely to Happen Next? Video-and-Language Future Event PredictionJie Lei, Licheng Yu, Tamara L. Berg, Mohit BansalEMNLP 2020 · 被引用 45 次
它引用的顶会 Paper5
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li 等ICLR 2020 · 被引用 1,825 次
- VaTeX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language ResearchXin Wang, Jiawei Wu, Jun-Kun Chen, Lei Li 等ICCV 2019 · 被引用 688 次
- CLEVRER: Collision Events for Video Representation and ReasoningKexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli 等ICLR 2020 · 被引用 584 次
- Relation-Aware Graph Attention Network for Visual Question AnsweringLinjie Li, Zhe Gan, Yu Cheng, Jingjing LiuICCV 2019 · 被引用 391 次
- TVQA+: Spatio-Temporal Grounding for Video Question AnsweringJie Lei, Licheng Yu, Tamara L. Berg, Mohit BansalACL 2020 · 被引用 173 次
相关 Paper
- Cross-modal Observation Hypothesis InferenceMengze Li, Kairong Han, Jiahe Xu, Yueying Li 等ACM MM 2024
- Adaptive Hierarchical Graph Reasoning with Semantic Coherence for Video-and-Language InferenceJuncheng Li, Siliang Tang, Linchao Zhu, Haochen Shi 等ICCV 2021 · 被引用 28 次
- SynopGround: A Large-Scale Dataset for Multi-Paragraph Video Grounding from TV Dramas and SynopsesChaolei Tan, Zihang Lin, Junfu Pu, Zhongang Qi 等ACM MM 2024 · 被引用 2 次
- ACQUIRED: A Dataset for Answering Counterfactual Questions In Real-Life VideosTe-Lin Wu, Zi-Yi Dou, Qingyuan Hu, Yu Hou 等EMNLP 2023 · 被引用 3 次
- InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and GenerationYi Wang, Yinan He, Yizhuo Li, Kunchang Li 等ICLR 2024 · 被引用 467 次
