UFVideo: Towards Unified Fine-Grained Video Cooperative Understanding with Large Language Models
Hewen Pan, Cong Wei, Dashuang Liang, Zepeng Huang, Pengfei Gao, Ziqi Zhou, Lulu Xue, Pengfei Yan, Xiaoming Wei, Minghui Li, Shengshan Hu
Abstract
A: This video depicts baseball players practicing throwing, and the presence of many teammates indicates that the sport heavily relies on teamwork.
Q: Why might object_1<region> be focused on dribbling two basketballs simultaneously?
A: Qobject_1 may be practicing his ball-handling skills to… Q: When does 'person opening a closet' happen in the video?
A: The event happen at 3.2s -7.3s.
Q: What color are the balance balls used by the athletes? Please generate a mask for the athlete on the right in each frame.
A: The balance balls used by athletes are blue. The mask: [SEG].
Q: What object_1: Man wearing black beanie eating food at table, object_2: Man in gray shirt focused on eating meal, object_3: Partial view of person's shoulder and arm at table is doing in the 0.7s, and generate the mask?
A: The video features object_1, a man wearing a black beanie...Object_2, another man in a gray shirt…Object_3, a partially visible person's shoulder and arm… Q: There are 3 objects in the video: object_1<region>, object_2<region>, object_3<region>. How does the position of object_2 change during the sequence? And please generate the mask in every frames?
A: Object_2 remains stationary beside object_1 throughout the sequence and does not move, even as object_3 kicks object_1.
Q: What object_1: Orange-handled hacksaw partially visible on right edge is doing, generate the time period and object mask?
A: The time is 3.2s -7.3s. From left to right, there are several traditional hand...To the right of this is a small, thin-bladed hacksaw with a metallic frame and red accents. Further right, an orange-handled hacksaw (object_1) is…
Referring Video object Segmentation Q: Please segment and track rabbit that gets food from the human.
A: The segmentation mask: [SEG]. Figure 1. Illustration of the video understanding tasks supported by UFVideo. To the best of our knowledge, UFVideo is the first unified method flexibly supports general video understanding, temporal video grounding, video object referring, referring and reasoning video object segmentation. Representative examples from a novel UFVideo-Bench demonstrating joint multi-grained video cooperative understanding through three tasks: PixRQA, PixHQA and PixTRQA.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 24616f67-cf66-4f52-9108-567905656ce3Builds on47
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- Frozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalMax Bain, Arsha Nagrani, Gül Varol, Andrew ZissermanICCV 2021 · 1,550 citations
- Grounding Multimodal Large Language Models to the WorldZhiliang Peng, Wenhui Wang, Li Dong, Yaru Hao et al.ICLR 2024 · 1,170 citations
Related papers
- SEED-Bench: Benchmarking Multimodal Large Language ModelsBohao Li, Yuying Ge, Yixiao Ge, Guangzhi Wang et al.CVPR 2024
- Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric VideosChiara Plizzari, Alessio Tonioni, Yongqin Xian, Achin Kulshrestha et al.CVPR 2025
- BASKET: A Large-Scale Video Dataset for Fine-Grained Skill EstimationYulu Pan, Ce Zhang, Gedas BertasiusCVPR 2025
- Action-Sketcher: From Reasoning to Action via Visual Sketches for Robotic ManipulationHuajie Tan, Peterson Co, Yijie Xu, Shanyu Rong et al.CVPR 2026
- Learning Asynchronous and Sparse Human-Object Interaction in VideosRomero F. A. B. de Morais, Vuong Le, Svetha Venkatesh, Truyen TranCVPR 2021
