ViCaS: A Dataset for Combining Holistic and Pixel-level Video Understanding using Captions with Grounded Segmentation
Ali Athar, Xueqing Deng, Liang-Chieh Chen
Abstract
Recent advances in multimodal large language models (MLLMs) have expanded research in video understanding, primarily focusing on high-level tasks such as video captioning and question-answering. Meanwhile, a smaller body of work addresses dense, pixel-precise segmentation tasks, which typically involve category-guided or referralbased object segmentation. Although both directions are essential for developing models with human-level video comprehension, they have largely evolved separately, with distinct benchmarks and architectures. This paper aims to unify these efforts by introducing ViCaS, a new dataset containing thousands of challenging videos, each annotated with detailed, human-written captions and temporally consistent, pixel-accurate masks for multiple objects with phrase grounding. Our benchmark evaluates models on both holistic/high-level understanding and languageguided, pixel-precise segmentation. We also present carefully validated evaluation measures and propose an effective model architecture that can tackle our benchmark.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f6a549e4-a798-4d11-9f73-c102568760caCited by top-tier papers5
- Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and GroundingChristopher Clark, Jieyu Zhang, Zixian Ma, Jae Sung Park et al.CVPR 2026 · 144 citations
- UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual ReasoningYe Liu, Zongyang Ma, Junfu Pu, Zhongang Qi et al.NeurIPS 2025 · 39 citations
- SPOT: Spatiotemporal Prompt Optimization for Motion-Stabilized MLLM-Guided Video SegmentationJiayi Fan, Zheyun Qin, Xiaoming Xi, Xiushan Nie et al.CVPR 2026
- From Words to Pixels: A Comprehensive Survey on Large Language Models in Visual SegmentationYizhou Wang, Mang Tik Chiu, Lingzhi Zhang, Xuan Shen et al.ACL 2026
- Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image CaptionsYijun Shen, Delong Chen, Fan Liu, Xingyu Wang et al.EMNLP 2025
Builds on50
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
Related papers
- Instruction-guided Multi-Granularity Segmentation and Captioning with Large Multimodal ModelXu Yuan, Li Zhou, Zenghui Sun, Zikun Zhou et al.AAAI 2025 · 1 citation
- SAMA: Towards Multi-Turn Referential Grounded Video Chat with Large Language ModelsYe Sun, Hao Zhang, Henghui Ding, Tiehua Zhang et al.NeurIPS 2025 · 9 citations
- VQAGuider: Guiding Multimodal Large Language Models to Answer Complex Video QuestionsYuyan Chen, Jiyuan Jia, Jiaxin Lu, Siyue Li et al.ACL 2025
- VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLMYuqian Yuan, Hang Zhang, Wentong Li, Zesen Cheng et al.CVPR 2025
- Object-Centric Video Question Answering with Visual Grounding and ReferringHaochen Wang, Qirui Chen, Cilin Yan, Jiayin Cai et al.ICCV 2025 · 2 citations
