Open-Vocabulary Video Relation Extraction
Wentao Tian, Zheng Wang, Yuqian Fu, Jingjing Chen, Lechao Cheng
摘要
A comprehensive understanding of videos is inseparable from describing the action with its contextual action-object interactions. However, many current video understanding tasks prioritize general action classification and overlook the actors and relationships that shape the nature of the action, resulting in a superficial understanding of the action. Motivated by this, we introduce Open-vocabulary Video Relation Extraction (OVRE), a novel task that views action understanding through the lens of action-centric relation triplets. OVRE focuses on pairwise relations that take part in the action and describes these relation triplets with natural languages. Moreover, we curate the Moments-OVRE dataset, which comprises 180K videos with action-centric relation triplets, sourced from a multi-label action classification dataset. With Moments-OVRE, we further propose a cross-modal mapping model to generate relation triplets as a sequence. Finally, we benchmark existing cross-modal generation models on the new task of OVRE. Our code and dataset are available at https://github.com/Iriya99/OVRE.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- VaTeX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language ResearchXin Wang, Jiawei Wu, Jun-Kun Chen, Lei Li 等ICCV 2019 · 被引用 688 次
- RegionCLIP: Region-based Language-Image PretrainingYiwu Zhong, Jianwei Yang, Pengchuan Zhang, Chunyuan Li 等CVPR 2022 · 被引用 481 次
相关 Paper
- Compositional Prompt Tuning with Motion Cues for Open-vocabulary Video Relation DetectionKaifeng Gao, Long Chen, Hanwang Zhang, Jun Xiao 等ICLR 2023 · 被引用 9 次
- Interventional Video Relation DetectionYicong Li, Xun Yang, Xindi Shang, Tat-Seng ChuaACM MM 2021 · 被引用 61 次
- Retrieval over Classification: Integrating Relation Semantics for Multimodal Relation ExtractionLei Hei, Tingjing Liao, Peiyingxin, Yiyang Qi 等EMNLP 2025
- Vision-Language Interactive Relation Mining for Open-Vocabulary Scene Graph GenerationYukuan Min, Muli Yang, Jinhao Zhang, Yuxuan Wang 等ICCV 2025 · 被引用 2 次
- Weakly Supervised Human-Object Interaction Detection in Video via Contrastive Spatiotemporal RegionsShuang Li, Yilun Du, Antonio Torralba, Josef Sivic 等ICCV 2021 · 被引用 17 次
