Tracking by 3D Model Estimation of Unknown Objects in Videos
Denys Rozumnyi, Jirí Matas, Marc Pollefeys, Vittorio Ferrari, Martin R. Oswald
Abstract
Most model-free visual object tracking methods formulate the tracking task as object location estimation given by a 2D segmentation or a bounding box in each video frame. We argue that this representation is limited and instead propose to guide and improve 2D tracking with an explicit object representation, namely the textured 3D shape and 6DoF pose in each video frame. Our representation tackles a complex long-term dense correspondence problem between all 3D points on the object for all video frames, including frames where some points are invisible. To achieve that, the estimation is driven by re-rendering the input video frames as well as possible through differentiable rendering, which has not been used for tracking before. The proposed optimization minimizes a novel loss function to estimate the best 3D shape, texture, and 6DoF pose. We improve the state-of-the-art in 2D segmentation tracking on three different datasets with mostly rigid objects.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 13e4b815-2ac1-41e2-9936-7fcae33f4951Cited by top-tier papers2
- PMGS: Reconstruction of Projectile Motion Across Large Spatiotemporal Spans via 3D Gaussian SplattingYijun Xu, Jingrui Zhang, Yuhan Chen, Dingwen Wang et al.AAAI 2026 · 1 citation
- Prior-free 3D Object TrackingXiuqiang Song, Li Jin, Zhengxian Zhang, Jiachen Li et al.CVPR 2025
Builds on13
- Pixel-Perfect Structure-from-Motion with Featuremetric RefinementPhilipp Lindenberger, Paul-Edouard Sarlin, Viktor Larsson, Marc PollefeysICCV 2021 · 266 citations
- Sharf: Shape-conditioned Radiance Fields from a Single ViewKonstantinos Rematas, Ricardo Martin-Brualla, Vittorio FerrariICML 2021 · 122 citations
- Multi-view 3D Reconstruction with TransformersDan Wang, Xinrui Cui, Xun Chen, Zhengxia Zou et al.ICCV 2021 · 111 citations
- CDTB: A Color and Depth Visual Object Tracking Dataset and BenchmarkAlan Lukezic, Ugur Kart, Jani Käpylä, Ahmed Durmush et al.ICCV 2019 · 79 citations
- Coupled Iterative Refinement for 6D Multi-Object Pose EstimationLahav Lipson, Zachary Teed, Ankit Goyal, Jia DengCVPR 2022 · 64 citations
Related papers
- GenFlow: Generalizable Recurrent Flow for 6D Pose Refinement of Novel ObjectsSungphill Moon, Hyeontae Son, Dongcheol Hur, Sangwook KimCVPR 2024 · 20 citations
- FoundationPose: Unified 6D Pose Estimation and Tracking of Novel ObjectsBowen Wen, Wei Yang, Jan Kautz, Stan BirchfieldCVPR 2024 · 215 citations
- 6DOPE-GS: Online 6D Object Pose Estimation using Gaussian SplattingYufeng Jin, Vignesh Prasad, Snehal Jauhri, Mathias Franzius et al.ICCV 2025 · 3 citations
- Shape-Constraint Recurrent Flow for 6D Object Pose EstimationYang Hai, Rui Song, Jiaojiao Li, Yinlin HuCVPR 2023
- Sub-Frame Appearance and 6D Pose Estimation of Fast Moving ObjectsDenys Rozumnyi, Jan Kotera, Filip Sroubek, Jiri MatasCVPR 2020
