SPOT: Spatiotemporal Prompt Optimization for Motion-Stabilized MLLM-Guided Video Segmentation
Jiayi Fan, Zheyun Qin, Xiaoming Xi, Xiushan Nie, Yilong Yin
摘要
The synergistic framework of multimodal large language models (MLLMs) and vision foundation models demonstrates exceptional performance in image understanding tasks, yet encounters severe temporal inconsistency challenges in video segmentation scenarios. Existing methods predominantly rely on MLLMs trained on static images to generate per-frame segmentation prompts, neglecting the physical continuity of video motion. This paper posits that performance limitations in video understanding tasks from inadequate constraints on model output behavior. Consequently, we propose a spatiotemporal co-optimization mechanism that achieves temporally consistent video segmentation solely by constraining MLLM output behavior, eliminating the need for large-scale video pretraining or complex architectural modifications. Our method features two complementary mechanisms: a Brownian bridge loss that models object trajectories as endpoint-constrained Gaussian processes to ensure temporal smoothness, and a geometry-aware prompt quality loss that enforces spatial consistency with target structures. Experiments on referring expression video segmentation and reasoning video segmentation tasks demonstrate that our method significantly surpasses state-of-the-art techniques on the Ref-YouTube-VOS, Ref-DAVIS-2017, MeVIS, A2d-Sentences, JHMDB-Sentences and ReVOS benchmarks. This work establishes that explicit modeling of physical world constraints can unlock the full potential of statically trained foundation models in dynamic visual understanding tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper23
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- MeViS: A Large-scale Benchmark for Video Segmentation with Motion ExpressionsHenghui Ding, Chang Liu, Shuting He, Xudong Jiang 等ICCV 2023 · 被引用 242 次
- Video-LLaVA: Learning United Visual Representation by Alignment Before ProjectionBin Lin, Yang Ye, Bin Zhu, Jiaxi Cui 等EMNLP 2024 · 被引用 231 次
- End-to-End Referring Video Object Segmentation with Multimodal TransformersAdam Botach, Evgenii Zheltonozhskii, Chaim BaskinCVPR 2022 · 被引用 150 次
- One Token to Seg Them All: Language Instructed Reasoning Segmentation in VideosZechen Bai, Tong He, Haiyang Mei, Pichao Wang 等NeurIPS 2024 · 被引用 147 次
相关 Paper
- Training-Free Spatio-temporal Decoupled Reasoning Video Segmentation with Adaptive Object MemoryZhengtong Zhu, Jiaqing Fan, Zhixuan Liu, Fanzhang LiAAAI 2026 · 被引用 1 次
- GLUS: Global-Local Reasoning Unified into A Single Large Language Model for Video SegmentationLang Lin, Xueyang Yu, Ziqi Pang, Yu-Xiong WangCVPR 2025
- DeRVOS: Decoupling Consistent Trajectory Generation and Multimodal Understanding for Referring Video Object SegmentationWenxuan Cheng, Ming Dai, Huimin Lu, Wankou YangCVPR 2026
- Object-Centric Video Question Answering with Visual Grounding and ReferringHaochen Wang, Qirui Chen, Cilin Yan, Jiayin Cai 等ICCV 2025 · 被引用 2 次
- ViLLa: Video Reasoning Segmentation with Large Language ModelRongkun Zheng, Lu Qi, Xi Chen, Yi Wang 等ICCV 2025 · 被引用 7 次
