UVLM: Benchmarking Video Language Model for Underwater World Understanding
Xizhe Xue, Yang Zhou, Dawei Yan, Lijie Tao, Junjie Li, Ying Li, Haokui Zhang, Rong Xiao
摘要
Recently, video-language models (VidLMs) have gained widespread attention and adoption. However, existing works primarily focus on terrestrial scenarios, overlooking the highly demanding application needs of underwater observation. To overcome this gap, we introduce UVLM, an under water observation benchmark which is build through a collaborative approach combining human expertise and AI models. To ensure data quality, we have conducted in-depth considerations from multiple perspectives. First, to address the unique challenges of underwater environments, we selected videos that represent typical underwater challenges including light variations, water turbidity, and diverse viewing angles to construct the dataset. Second, to ensure data diversity, the dataset covers a wide range of frame rates, resolutions, 419 classes of marine animals, and various static plants and terrains. Next, for task diversity, we adopted a structured design where observation targets are categorized into two major classes: biological and environmental. Each category includes content observation and change/action observation, totaling 20 subtask types. Finally, we designed several challenging evaluation metrics to enable quantitative comparison and analysis of different methods. Experiments on two representative VidLMs demonstrate that fine-tuning VidLMs on UVLM significantly improves underwater world understanding while also showing potential for slight improvements on existing in-air VidLM benchmarks, such as VideoMME and Perception text. The dataset and prompt are publicly available at: https://github.com/Cecilia-xue/UVLM-Benchmark .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper11
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi 等ICCV 2019 · 被引用 1,437 次
- InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and GenerationYi Wang, Yinan He, Yizhuo Li, Kunchang Li 等ICLR 2024 · 被引用 467 次
- MovieChat: From Dense Token to Sparse Memory for Long Video UnderstandingEnxin Song, Wenhao Chai, Guanhong Wang, Yucheng Zhang 等CVPR 2024 · 被引用 95 次
- WaterMask: Instance Segmentation for Underwater ImageryShijie Lian, Hua Li, Runmin Cong, Suqi Li 等ICCV 2023 · 被引用 72 次
- Self-supervised Monocular Underwater Depth Recovery, Image Restoration, and a Real-sea Video DatasetNisha Varghese, Ashish Kumar, A. N. RajagopalanICCV 2023 · 被引用 37 次
相关 Paper
- UQ-Bench: A Benchmark for Evaluating Multimodal LLMs on Underwater Image Quality AssessmentJingchao Cao, Guo An, Feng Gao, Ke Gu 等AAAI 2026
- Exploring the Underwater World Segmentation without Extra TrainingBingyu Li, Tao Huo, Da Zhang, Zhiyuan Zhao 等CVPR 2026 · 被引用 18 次
- Is your VLM Sky-Ready? A Comprehensive Spatial Intelligence Benchmark for UAV NavigationLingfeng Zhang, Yuchen Zhang, Hongsheng Li, Haoxiang Fu 等CVPR 2026 · 被引用 15 次
- NAUTILUS: A Large Multimodal Model for Underwater Scene UnderstandingWei Xu, Cheng Wang, Dingkang Liang, Zongchuang Zhao 等NeurIPS 2025 · 被引用 16 次
- Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video AnalysisChaoyou Fu, Yuhan Dai, Yongdong Luo, Lei Li 等CVPR 2025
