RSPT: Reconstruct Surroundings and Predict Trajectory for Generalizable Active Object Tracking
Fangwei Zhong, Xiao Bi, Yudi Zhang, Wei Zhang, Yizhou Wang
Abstract
Active Object Tracking (AOT) aims to maintain a specific relation between the tracker and object(s) by autonomously controlling the motion system of a tracker given observations. AOT has wide-ranging applications, such as in mobile robots and autonomous driving. However, building a generalizable active tracker that works robustly across different scenarios remains a challenge, especially in unstructured environments with cluttered obstacles and diverse layouts. We argue that constructing a state representation capable of modeling the geometry structure of the surroundings and the dynamics of the target is crucial for achieving this goal. To address this challenge, we present RSPT, a framework that forms a structure-aware motion representation by Reconstructing the Surroundings and Predicting the target Trajectory. Additionally, we enhance the generalization of the policy network by training in an asymmetric dueling mechanism. We evaluate RSPT on various simulated scenarios and show that it outperforms existing methods in unseen environments, particularly those with complex obstacles and layouts. We also demonstrate the successful transfer of RSPT to real-world settings. .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 24b0eac2-3ae6-43cd-b8f5-66a7a46ed891Cited by top-tier papers5
- Richelieu: Self-Evolving LLM-Based Agents for AI DiplomacyZhenyu Guan, Xiangyu Kong, Fangwei Zhong, Yizhou WangNeurIPS 2024 · 48 citations
- Fast Peer Adaptation with Context-aware ExplorationLong Ma, Yuanfei Wang, Fangwei Zhong, Song-Chun Zhu et al.ICML 2024 · 9 citations
- Multimodal Sense-Informed Forecasting of 3D Human MotionsZhenyu Lou, Qiongjie Cui, Haofan Wang, Xu Tang et al.CVPR 2024 · 8 citations
- UnrealZoo: Enriching Photo-Realistic Virtual Worlds for Embodied AIFangwei Zhong, Kui Wu, Churan Wang, Hao Chen et al.ICCV 2025 · 6 citations
- Instance-level Visual Active Tracking with Occlusion-Aware PlanningHaowei Sun, Kai Zhou, Hao Gao, Shiteng Zhang et al.CVPR 2026 · 4 citations
Builds on6
- Learning Discriminative Model Prediction for TrackingGoutam Bhat, Martin Danelljan, Luc Van Gool, Radu TimofteICCV 2019 · 1,294 citations
- ToM2C: Target-oriented Multi-agent Communication and Cooperation with Theory of MindYuanfei Wang, Fangwei Zhong, Jing Xu, Yizhou WangICLR 2022 · 103 citations
- Learning Multi-Agent Coordination for Enhancing Target Coverage in Directional Sensor NetworksJing Xu, Fangwei Zhong, Yizhou WangNeurIPS 2020 · 71 citations
- Pose-Assisted Multi-Camera Collaboration for Active Object TrackingJing Li, Jing Xu, Fangwei Zhong, Xiangyu Kong et al.AAAI 2020 · 55 citations
- Towards Distraction-Robust Active Visual TrackingFangwei Zhong, Peng Sun, Wenhan Luo, Tingyun Yan et al.ICML 2021 · 50 citations
Related papers
- UAST: Unified Active Search and Tracking for Arbitrary Targets with UAVsLiang Qin, Min Wang, Xingyu Lu, Aowen Qiu et al.CVPR 2026
- Paparazzo: Active Mapping of Moving 3D ObjectsDavide Allegro, Shiyao Li, Stefano Ghidoni, Vincent LepetitCVPR 2026
- Generalizable Structure-Aware Keypoint Correspondence for Category-Unified 3D Single Object TrackingJie Xiao, Yinchao Ma, Yuyang Tang, Dengqing Yang et al.CVPR 2026
- Motion-Aware Object Tracking via Motion and Geometry-Aware CuesHongtao Yang, Bineng Zhong, Qihua Liang, Xiantao Hu et al.AAAI 2026
- Visual Representation Learning with Stochastic Frame PredictionHuiwon Jang, Dongyoung Kim, Junsu Kim, Jinwoo Shin et al.ICML 2024 · 10 citations
