DTOS: Dynamic Time Object Sensing with Large Multimodal Model
Jirui Tian, Jinrong Zhang, Shenglan Liu, Luhao Xu, Zhixiong Huang, Gao Huang
摘要
Existing multimodal large language models (MLLMs) face significant challenges in Referring Video Object Segmentation(RVOS). We identify three critical challenges: (C1) insufficient quantitative representation of textual numerical data, (C2) repetitive and degraded response templates for spatiotemporal referencing, and (C3) loss of visual information in video sampling queries lacking textual guidance. To address these, we propose a novel framework, Dynamic Time Object Sensing (DTOS), specifically designed for RVOS. To tackle (C1) and (C2), we introduce specialized tokens to construct multi-answer response templates, enabling regression of event boundaries and target localization. This approach improves the accuracy of numerical regression while mitigating the issue of repetitive degradation. To address (C3), we propose a Text-guided Clip Sampler (TCS) that selects video clips aligned with user instructions, preventing visual information loss and ensuring consistent temporal resolution. TCS is also applicable to Moment Retrieval tasks, with enhanced multimodal input sequences preserving spatial details and maximizing temporal resolution. DTOS demonstrates exceptional capability in flexibly localizing multiple spatiotemporal targets based on userprovided textual instructions. Extensive experiments validate the effectiveness of our approach, with DTOS achieving state-of-the-art performance in
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- SPOT: Spatiotemporal Prompt Optimization for Motion-Stabilized MLLM-Guided Video SegmentationJiayi Fan, Zheyun Qin, Xiaoming Xi, Xiushan Nie 等CVPR 2026
- Breaking the Regional Perception Bottleneck of Multimodal Large Language Models via External Reasoning FrameworkJinrong Zhang, Zhaoyang Xu, Xusheng He, Xinrui Li 等CVPR 2026
它引用的顶会 Paper28
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- Detecting Moments and Highlights in Videos via Natural Language QueriesJie Lei, Tamara L. Berg, Mohit BansalNeurIPS 2021 · 被引用 425 次
- Self-Chained Image-Language Model for Video Localization and Question AnsweringShoubin Yu, Jaemin Cho, Prateek Yadav, Mohit BansalNeurIPS 2023 · 被引用 281 次
- Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language ModelsMuhammad Maaz, Hanoona Abdul Rasheed, Salman Khan, Fahad KhanACL 2024 · 被引用 279 次
相关 Paper
- The Devil is in Temporal Token: High Quality Video Reasoning SegmentationSitong Gong, Yunzhi Zhuge, Lu Zhang, Zongxin Yang 等CVPR 2025
- VideoOrion: Tokenizing Object Dynamics in VideosYicheng Feng, Yijiang Li, Wanpeng Zhang, Sipeng Zheng 等ICCV 2025
- GroundVTS: Visual Token Sampling in Multimodal Large Language Models for Video Temporal GroundingRong Fan, Kaiyan Xiao, Minghao Zhu, Liuyi Wang 等CVPR 2026 · 被引用 1 次
- GLUS: Global-Local Reasoning Unified into A Single Large Language Model for Video SegmentationLang Lin, Xueyang Yu, Ziqi Pang, Yu-Xiong WangCVPR 2025
- MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object SegmentationFu Rong, Meng Lan, Qian Zhang, Lefei ZhangICCV 2025 · 被引用 4 次
