Lune

CVPR2026顶会

UFVideo: Towards Unified Fine-Grained Video Cooperative Understanding with Large Language Models

Hewen Pan, Cong Wei, Dashuang Liang, Zepeng Huang, Pengfei Gao, Ziqi Zhou, Lulu Xue, Pengfei Yan, Xiaoming Wei, Minghui Li, Shengshan Hu

2026年份
3被引次数

摘要

A: This video depicts baseball players practicing throwing, and the presence of many teammates indicates that the sport heavily relies on teamwork.

Q: Why might object_1<region> be focused on dribbling two basketballs simultaneously?

A: Qobject_1 may be practicing his ball-handling skills to… Q: When does 'person opening a closet' happen in the video?

A: The event happen at 3.2s -7.3s.

Q: What color are the balance balls used by the athletes? Please generate a mask for the athlete on the right in each frame.

A: The balance balls used by athletes are blue. The mask: [SEG].

Q: What object_1: Man wearing black beanie eating food at table, object_2: Man in gray shirt focused on eating meal, object_3: Partial view of person's shoulder and arm at table is doing in the 0.7s, and generate the mask?

A: The video features object_1, a man wearing a black beanie...Object_2, another man in a gray shirt…Object_3, a partially visible person's shoulder and arm… Q: There are 3 objects in the video: object_1<region>, object_2<region>, object_3<region>. How does the position of object_2 change during the sequence? And please generate the mask in every frames?

A: Object_2 remains stationary beside object_1 throughout the sequence and does not move, even as object_3 kicks object_1.

Q: What object_1: Orange-handled hacksaw partially visible on right edge is doing, generate the time period and object mask?

A: The time is 3.2s -7.3s. From left to right, there are several traditional hand...To the right of this is a small, thin-bladed hacksaw with a metallic frame and red accents. Further right, an orange-handled hacksaw (object_1) is…

Referring Video object Segmentation Q: Please segment and track rabbit that gets food from the human.

A: The segmentation mask: [SEG]. Figure 1. Illustration of the video understanding tasks supported by UFVideo. To the best of our knowledge, UFVideo is the first unified method flexibly supports general video understanding, temporal video grounding, video object referring, referring and reasoning video object segmentation. Representative examples from a novel UFVideo-Bench demonstrating joint multi-grained video cooperative understanding through three tasks: PixRQA, PixHQA and PixTRQA.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper47

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖