VideoComp: Advancing Fine-Grained Compositional and Temporal Alignment in Video-Text Models
Dahun Kim, A. J. Piergiovanni, Ganesh Satish Mallya, Anelia Angelova
Abstract
We introduce VideoComp, a benchmark and learning framework for advancing video-text compositionality understanding, aimed at improving vision-language models (VLMs) in fine-grained temporal alignment. Unlike existing benchmarks focused on static image-text compositionality or isolated single-event videos, our benchmark targets alignment in continuous multi-event videos. Leveraging videotext datasets with temporally localized event captions (e.g. ActivityNet-Captions, YouCook2), we construct two compositional benchmarks, ActivityNet-Comp and YouCook2-Comp. We create challenging negative samples with subtle temporal disruptions such as reordering, action word replacement, partial captioning, and combined disruptions. These benchmarks comprehensively test models' compositional sensitivity across extended, cohesive video-text sequences. To improve model performance, we propose a hierarchical pairwise preference loss that strengthens alignment with temporally accurate pairs and gradually penalizes increasingly disrupted ones, encouraging fine-grained compositional learning. To mitigate the limited availability of densely annotated video data, we introduce a pretraining strategy that concatenates short video-caption pairs to simulate multi-event sequences. We evaluate video-text foundational models and large multimodal models (LMMs) on our benchmark, identifying both strengths and areas for improvement in compositionality. Overall, our work provides a comprehensive framework for evaluating and enhancing model capabilities in achieving fine-grained, temporally coherent video-text alignment. Dataset available at: https://github.com/google-deepmind/video comp.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d2033179-8d9b-490f-a278-1d1aece4bf11Cited by top-tier papers3
- Dynamic Reflections: Probing Video Representations with Text AlignmentMaks Ovsjanikov, Viorica Patraucean, Leonidas J. Guibas, Tyler Zhu et al.ICLR 2026 · 5 citations
- SVHalluc: Benchmarking Speech-Vision Hallucination in Audio-Visual Large Language ModelsChenshuang Zhang, Kyeong Seon Kim, Chengxin Liu, Tae-Hyun OhCVPR 2026 · 2 citations
- What "Not" to Detect: Negation-Aware VLMs via Structured Reasoning and Token MergingInha Kang, Youngsun Lim, Seonho Lee, Jiho Choi et al.ICLR 2026 · 1 citation
Builds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Frozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalMax Bain, Arsha Nagrani, Gül Varol, Andrew ZissermanICCV 2021 · 1,550 citations
- InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and GenerationYi Wang, Yinan He, Yizhuo Li, Kunchang Li et al.ICLR 2024 · 467 citations
- VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text UnderstandingHu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko et al.EMNLP 2021 · 399 citations
Related papers
- VidComposition: Can MLLMs Analyze Compositions in Compiled Videos?Yunlong Tang, Junjia Guo, Hang Hua, Susan Liang et al.CVPR 2025
- VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality EvaluationShi-Xue Zhang, Hongfa Wang, Duojun Huang, Xin Li et al.AAAI 2026 · 5 citations
- Structured Video-Language Modeling with Temporal Grouping and Spatial GroundingYuanhao Xiong, Long Zhao, Boqing Gong, Ming-Hsuan Yang et al.ICLR 2024
- DisTime: Distribution-Based Time Representation for Video Large Language ModelsYingsen Zeng, Zepeng Huang, Yujie Zhong, Chengjian Feng et al.ICCV 2025 · 2 citations
- Compositional Temporal Grounding with Structured Variational Cross-Graph Correspondence LearningJuncheng Li, Junlin Xie, Long Qian, Linchao Zhu et al.CVPR 2022 · 63 citations
