LoVR: A Benchmark for Long Video Retrieval in Multimodal Contexts
Hao Liang, Qifeng Cai, Zhaoyang Han, Hejun Dong, Meiyi Qiang, Ruichuan An, Quanqing Xu, Bin Cui, Wentao Zhang
摘要
Long videos contain a vast amount of information, making videotext retrieval an essential and challenging task in multimodal learning and web-scale search. On today's Web, where users increasingly expect to locate not only relevant pages but also specific long videos or fine-grained clips, existing benchmarks fall short due to limited video duration, low-quality captions, and coarse annotation granularity. To address these limitations, we introduce LoVR, a benchmark specifically designed for long video-text retrieval. LoVR contains 467 long videos and over 40,804 fine-grained clips with high-quality captions. To overcome the issue of poor machinegenerated annotations, we propose an efficient caption generation framework that integrates VLM automatic generation, caption quality scoring, and dynamic refinement. This pipeline improves annotation accuracy while maintaining scalability. Furthermore, we introduce a semantic fusion method to generate coherent fullvideo captions without losing important contextual information. Our benchmark introduces longer videos, more detailed captions, and a larger-scale dataset, presenting new challenges for video understanding and retrieval. Extensive experiments on various advanced models demonstrate that LoVR is a challenging benchmark, revealing the limitations of current approaches and providing valuable insights for future research. We release the code link at https://lovrbench.github.io/ * Equal contribution. †Corresponding author.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- WAVE: Learning Unified & Versatile Audio-Visual Embeddings with Multimodal LLMChangli Tang, Qinfan Xiao, Ke Mei, Tianyi Wang 等ICLR 2026 · 被引用 9 次
- RecTok: Reconstruction Distillation along Rectified FlowQingyu Shi, Size Wu, Jinbin Bai, Kaidong Yu 等CVPR 2026 · 被引用 5 次
它引用的顶会 Paper12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi 等ICCV 2019 · 被引用 1,437 次
- LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic AlignmentBin Zhu, Bin Lin, Munan Ning, Yang Yan 等ICLR 2024 · 被引用 403 次
- Perception Encoder: The best visual embeddings are not at the output of the networkDaniel Bolya, Po-Yao Huang, Peize Sun, Jang Hyun Cho 等NeurIPS 2025 · 被引用 359 次
相关 Paper
- VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality EvaluationShi-Xue Zhang, Hongfa Wang, Duojun Huang, Xin Li 等AAAI 2026 · 被引用 5 次
- ALLVB: All-in-One Long Video Understanding BenchmarkXichen Tan, Yuanjing Luo, Yunfan Ye, Fang Liu 等AAAI 2025 · 被引用 13 次
- CaReBench: A Fine-grained Benchmark for Video Captioning and RetrievalYifan Xu, Xinhao Li, Yichun Yang, Desen Meng 等ICLR 2026 · 被引用 10 次
- VISTA: Enhancing Long-Duration and High-Resolution Video Understanding by Video Spatiotemporal AugmentationWeiming Ren, Huan Yang, Jie Min, Cong Wei 等CVPR 2025
- AuroraCap: Efficient, Performant Video Detailed Captioning and a New BenchmarkWenhao Chai, Enxin Song, Yilun Du, Chenlin Meng 等ICLR 2025
