VSTAR: A Video-grounded Dialogue Dataset for Situated Semantic Understanding with Scene and Topic Transitions
Yuxuan Wang, Zilong Zheng, Xueliang Zhao, Jinpeng Li, Yueqian Wang, Dongyan Zhao
摘要
Video-grounded dialogue understanding is a challenging problem that requires machine to perceive, parse and reason over situated semantics extracted from weakly aligned video and dialogues. Most existing benchmarks treat both modalities the same as a frame-independent visual understanding task, while neglecting the intrinsic attributes in multimodal dialogues, such as scene and topic transitions. In this paper, we present Video-grounded Scene&Topic AwaRe dialogue (VSTAR) dataset, a large scale video-grounded dialogue understanding dataset based on 395 TV series. Based on VSTAR, we propose two benchmarks for video-grounded dialogue understanding: scene segmentation and topic segmentation, and one benchmark for video-grounded dialogue generation. Comprehensive experiments are performed on these benchmarks to demonstrate the importance of multimodal information and segments in video-grounded dialogue understanding and generation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Zebra-CoT: A Dataset for Interleaved Vision-Language ReasoningAng Li, Charles L. Wang, Deqing Fu, Kaiyu Yue 等ICLR 2026 · 被引用 85 次
- VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory BridgesYuxuan Wang, Yiqi Song, Cihang Xie, Yang Liu 等ICCV 2025 · 被引用 7 次
- Friends-MMC: A Dataset for Multi-modal Multi-party Conversation UnderstandingYueqian Wang, Xiaojun Meng, Yuxuan Wang, Jianxin Liang 等AAAI 2025 · 被引用 6 次
- STAIR: Spatial-Temporal Reasoning with Auditable Intermediate Results for Video Question AnsweringYueqian Wang, Yuxuan Wang, Kai Chen, Dongyan ZhaoAAAI 2024 · 被引用 4 次
- A Video-grounded Dialogue Dataset and Metric for Event-driven ActivitiesWiradee Imrattanatrai, Masaki Asada, Kimihiro Hasegawa, Zhi-Qi Cheng 等AAAI 2025
它引用的顶会 Paper6
- What Makes Multi-Modal Learning Better than Single (Provably)Yu Huang, Chenzhuang Du, Zihui Xue, Xuanyao Chen 等NeurIPS 2021 · 被引用 404 次
- Multimodal Dialogue Response GenerationQingfeng Sun, Yujing Wang, Can Xu, Kai Zheng 等ACL 2022 · 被引用 58 次
- A Local-to-Global Approach to Multi-Modal Movie Scene SegmentationAnyi Rao, Linning Xu, Yu Xiong, Guodong Xu 等CVPR 2020
- Shot Contrastive Self-Supervised Learning for Scene Boundary DetectionShixing Chen, Xiaohan Nie, David Fan, Dongqing Zhang 等CVPR 2021
- DVD: A Diagnostic Dataset for Multi-step Reasoning in Video Grounded DialogueHung Le, Chinnadhurai Sankar, Seungwhan Moon, Ahmad Beirami 等ACL 2021
相关 Paper
- SuperDialseg: A Large-scale Dataset for Supervised Dialogue SegmentationJunfeng Jiang, Chengzhang Dong, Sadao Kurohashi, Akiko AizawaEMNLP 2023 · 被引用 2 次
- SynopGround: A Large-Scale Dataset for Multi-Paragraph Video Grounding from TV Dramas and SynopsesChaolei Tan, Zihang Lin, Junfu Pu, Zhongang Qi 等ACM MM 2024 · 被引用 2 次
- Structured Co-reference Graph Attention for Video-grounded DialogueJunyeong Kim, Sunjae Yoon, Dahyun Kim, Chang D. YooAAAI 2021 · 被引用 31 次
- SAMA: Towards Multi-Turn Referential Grounded Video Chat with Large Language ModelsYe Sun, Hao Zhang, Henghui Ding, Tiehua Zhang 等NeurIPS 2025 · 被引用 9 次
- Violin: A Large-Scale Dataset for Video-and-Language InferenceJingzhou Liu, Wenhu Chen, Yu Cheng, Zhe Gan 等CVPR 2020
