OwlCap: Harmonizing Motion-Detail for Video Captioning via HMD-270K and Caption Set Equivalence Reward
Chunlin Zhong, Qiuxia Hou, Zhangjun Zhou, Yanhao Zhang, Shuang Hao, Haonan Lu, He Tang, Xiang Bai
摘要
Video captioning aims to generate comprehensive and coherent descriptions of the video content, contributing to the advancement of both video understanding and generation. However, existing methods often suffer from motion-detail imbalance, as models tend to overemphasize one aspect while neglecting the other. This imbalance results in incomplete captions, which in turn leads to a lack of consistency in video understanding and generation. To address this issue, we propose solutions from two aspects: 1) Data aspect: we constructed the Harmonizing Motion-Detail 270K (HMD-270K) dataset through a two-stage pipeline: Motion-Detail Fusion (MDF) and Fine-Grained Examination (FGE). 2) Optimization aspect: We introduce the Caption Set Equivalence Reward (CSER) based on Group Relative Policy Optimization (GRPO). CSER enhances completeness and accuracy in capturing both motion and details through unit-to-set matching and bidirectional validation. Based on the HMD-270K supervised fine-tuning and GRPO post-training with CSER, we developed OwlCap, a powerful video captioning multimodal large language model (MLLM) with motion-detail balance. Experimental results demonstrate that OwlCap achieves significant improvements compared to baseline models on two benchmarks: the detail-focused VDC (+4.2 Acc) and the motion-focused DREAM-1K (+4.6 F1). The HMD-270K dataset and OwlCap model will be publicly released to facilitate video captioning research community advancements.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- AVoCaDO: An Audiovisual Video Captioner Driven by Temporal OrchestrationXinlong Chen, Yue Ding, Weihong Lin, Jingyun Hua 等ICLR 2026 · 被引用 27 次
- Building a Precise Video Language with Human–AI OversightZhiqiu Lin, Siyuan Cen, Chancharik Mitra, Isaac Li 等CVPR 2026 · 被引用 3 次
- TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual CaptionsLinli Yao, Yuancheng Wei, Yaojie Zhang, Lei Li 等ICML 2026
它引用的顶会 Paper11
- Frozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalMax Bain, Arsha Nagrani, Gül Varol, Andrew ZissermanICCV 2021 · 被引用 1,550 次
- Visual-RFT: Visual Reinforcement Fine-TuningZiyu Liu, Zeyi Sun, Yuhang Zang, Xiaoyi Dong 等ICCV 2025 · 被引用 563 次
- Video-R1: Reinforcing Video Reasoning in MLLMsKaituo Feng, Kaixiong Gong, Bohao Li, Zonghao Guo 等NeurIPS 2025 · 被引用 528 次
- StableVideo: Text-driven Consistency-aware Diffusion Video EditingWenhao Chai, Xun Guo, Gaoang Wang, Yan LuICCV 2023 · 被引用 219 次
- CaReBench: A Fine-grained Benchmark for Video Captioning and RetrievalYifan Xu, Xinhao Li, Yichun Yang, Desen Meng 等ICLR 2026 · 被引用 10 次
相关 Paper
- AuroraCap: Efficient, Performant Video Detailed Captioning and a New BenchmarkWenhao Chai, Enxin Song, Yilun Du, Chenlin Meng 等ICLR 2025
- Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree SearchLinhao Yu, Xingguang Ji, Yahui Liu, Fanheng Kong 等ACL 2025 · 被引用 2 次
- Progress-Aware Video Frame CaptioningZihui Xue, Joungbin An, Xitong Yang, Kristen GraumanCVPR 2025
- LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language ModelsShenghao Fu, Qize Yang, Qijie Mo, Junkai Yan 等CVPR 2025
- Towards Fine-Grained Human Motion Video CaptioningGuorui Song, Guocun Wang, Zhe Huang, Jing Lin 等ACM MM 2025
