RealVG: Unleashing MLLMs for Training-Free Spatio-Temporal Video Grounding in the Wild
Hongchen Wei, Zhenzhong Chen
摘要
Spatio-Temporal Video Grounding (STVG) aims to localize spatio-temporal tubes of specific objects or actions within videos based on textual queries. Despite significant progress, existing methods struggle to generalize effectively to real-world scenarios due to the limited quantity and diversity of annotated data. In this paper, we introduce RealVG, a robust and training-free pipeline that leverages powerful Multimodal Large Language Models (MLLMs) through question-answering to tackle STVG in the wild. To address the challenges posed by complex real-world videos and queries, we propose a spatio-temporal decoupling module and a query-guided visual token filter to decompose intricate scenes and refine target-oriented perception, enhancing the robustness and adaptability of MLLMs. Specifically, the spatio-temporal decoupling module breaks down videos and queries into simpler sub-scenes and sub-queries, reducing complexity and promoting a precise understanding of static visual elements. Meanwhile, the query-guided visual token filter eliminates irrelevant tokens, sharpening focus on the target object and improving short-range action perception. Experimental results demonstrate that RealVG achieves superior performance over state-of-the-art supervised and weakly supervised methods in real-world settings, despite requiring no STVG data for training.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- Agentic Spatio-Temporal Grounding via Collaborative ReasoningHeng Zhao, Yew-Soon Ong, Joey Tianyi ZhouSIGIR 2026 · 被引用 1 次
- SARL-STG: A Spatially Aware Reinforcement Learning Framework for Refining MLLMs in Spatio-Temporal Video GroundingHong Gao, Xiangkai Xu, Bin Zhong, Junjie Yin 等CVPR 2026
相关 Paper
- Unleashing the Potential of Multimodal LLMs for Zero-Shot Spatio-Temporal Video GroundingZaiquan Yang, Yuhao Liu, Gerhard P. Hancke, Rynson W. H. LauNeurIPS 2025 · 被引用 10 次
- SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding CapabilityJiankang Wang, Zhihan Zhang, Zhihang Liu, Yang Li 等AAAI 2026 · 被引用 20 次
- OmniGround: A Comprehensive Spatio-Temporal Grounding Benchmark for Real-World Complex ScenariosHong Gao, Jingyu Wu, Xiangkai Xu, Kangni Xie 等CVPR 2026 · 被引用 4 次
- Universal Video Temporal Grounding with Generative Multi-modal Large Language ModelsZeqian Li, Shangzhe Di, Zhonghua Zhai, Weilin Huang 等NeurIPS 2025 · 被引用 30 次
- STVGBert: A Visual-linguistic Transformer based Framework for Spatio-temporal Video GroundingRui Su, Qian Yu, Dong XuICCV 2021 · 被引用 75 次
