CTO-SLAM: Contour Tracking for Object-Level Robust 4D SLAM
Xiaohan Li, Dong Liu, Jun Wu
Abstract
The demand for 4D ( 3D+time ) SLAM system is increasingly urgent, especially for decision-making and scene understanding. However, most of the existing simultaneous localization and mapping ( SLAM ) systems primarily assume static environments. They fail to represent dynamic scenarios due to the challenge of establishing robust long-term spatiotemporal associations in dynamic object tracking. We address this limitation and propose CTO-SLAM, a monocular and RGB-D object-level 4D SLAM system to track moving objects and estimate their motion simultaneously. In this paper, we propose contour tracking, which introduces contour features to enhance the keypoint representation of dynamic objects and coupled with pixel tracking to achieve long-term robust object tracking. Based on contour tracking, we propose a novel sampling-based object pose initialization algorithm and the following adapted bundle adjustment ( BA ) optimization algorithm to estimate dynamic object poses with high accuracy. The CTO-SLAM system is verified on both KITTI and VKITTI datasets. The experimental results demonstrate that our system effectively addresses cumulative errors in long-term spatiotemporal association and hence obtains substantial improvements over the state-of-the-art systems. The source code is available at https://github.com/realXiaohan/CTO-SLAM.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ca049762-2fcd-4246-9a16-9e54922b5479Cited by top-tier papers3
- Ov3R: Open-Vocabulary Semantic 3D Reconstruction from RGB VideosZiren Gong, Xiaohan Li, Fabio Tosi, Jiawei Han et al.CVPR 2026 · 13 citations
- Image-to-Point Cloud Feature Back-Projection for Multimodal Training of 3D Semantic SegmentationJiawei Han, Matteo Poggi, Li Huan, Changshuo Wang et al.CVPR 2026
- SAG-GNN: Semantic-Aware Guided GNN for Descriptor-Free 2D-3D MatchingShihua Zhang, Tianhao Xu, Zizhuo Li, Qing Ma et al.CVPR 2026
Builds on6
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D CamerasZachary Teed, Jia DengNeurIPS 2021 · 1,248 citations
- COTR: Correspondence Transformer for Matching Across ImagesWei Jiang, Eduard Trulls, Jan Hosang, Andrea Tagliasacchi et al.ICCV 2021 · 318 citations
- Towards An End-to-End Framework for Flow-Guided Video InpaintingZhen Li, Chengze Lu, Jianhua Qin, Chun-Le Guo et al.CVPR 2022 · 136 citations
- ClusterSLAM: A SLAM Backend for Simultaneous Rigid Body Clustering and Motion EstimationJiahui Huang, Sheng Yang, Zishuo Zhao, Yu-Kun Lai et al.ICCV 2019 · 89 citations
- Learning Fine-Grained Features for Pixel-wise Video CorrespondencesRui Li, Shenglong Zhou, Dong LiuICCV 2023 · 7 citations
Related papers
- DROID-SLAM in the WildMoyang Li, Zihan Zhu, Marc Pollefeys, Daniel BarathCVPR 2026 · 10 citations
- EM-Fusion: Dynamic Object-Level SLAM With Probabilistic Data AssociationMichael Strecke, Jörg StücklerICCV 2019 · 86 citations
- 4DTAM: Non-Rigid Tracking and Mapping via Dynamic Surface GaussiansHidenobu Matsuki, Gwangbin Bae, Andrew J. DavisonCVPR 2025
- GO-SLAM: Global Optimization for Consistent 3D Instant ReconstructionYoumin Zhang, Fabio Tosi, Stefano Mattoccia, Matteo PoggiICCV 2023 · 208 citations
- You Don't Only Look Once: Constructing Spatial-Temporal Memory for Integrated 3D Object Detection and TrackingJiaming Sun, Yiming Xie, Siyu Zhang, Linghao Chen et al.ICCV 2021 · 12 citations
