KPA-Tracker: Towards Robust and Real-Time Category-Level Articulated Object 6D Pose Tracking
Liu Liu, Anran Huang, Qi Wu, Dan Guo, Xun Yang, Meng Wang
Abstract
Our life is populated with articulated objects. Current category-level articulation estimation works largely focus on predicting part-level 6D poses on static point cloud observations. In this paper, we tackle the problem of category-level online robust and real-time 6D pose tracking of articulated objects, where we propose KPA-Tracker, a novel 3D KeyPoint based Articulated object pose Tracker. Given an RGB-D image or a partial point cloud at the current frame as well as the estimated per-part 6D poses from the last frame, our KPA-Tracker can effectively update the poses with learned 3D keypoints between the adjacent frames. Specifically, we first canonicalize the input point cloud and formulate the pose tracking as an inter-frame pose increment estimation task. To learn consistent and separate 3D keypoints for every rigid part, we build KPA-Gen that outputs the high-quality ordered 3D keypoints in an unsupervised manner. During pose tracking on the whole video, we further propose a keypoint-based articulation tracking algorithm that mines keyframes as reference for accurate pose updating. We provide extensive experiments on validating our KPA-Tracker on various datasets ranging from synthetic point cloud observation to real-world scenarios, which demonstrates the superior performance and robustness of the KPA-Tracker. We believe that our work has the potential to be applied in many fields including robotics, embodied intelligence and augmented reality. All the datasets and codes are available at https://github.com/hhhhhar/KPA-Tracker.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 09874d76-4a71-415a-92ef-0d2333147d3fCited by top-tier papers3
- AffordBot: 3D Fine-grained Embodied Reasoning via Multimodal Large Language ModelsXinyi Wang, Xun Yang, Yanlong Xu, Yuchen Wu et al.NeurIPS 2025 · 18 citations
- DICArt: Advancing Category-level Articulated Object Pose Estimation in Discrete State-SpacesLi Zhang, Mingyu Mei, Ailing Wang, Xianhui Meng et al.CVPR 2026 · 2 citations
- DFGAP: Towards Depth-Free Cross-Category GAParts Perception via Uncertainty-Quantified ModelingXueyu Yuan, Jiarui Zhang, Jiangqi Song, Liu Liu et al.ACM MM 2025
Builds on10
- GPV-Pose: Category-level Object Pose Estimation via Geometry-guided Point-wise VotingYan Di, Ruida Zhang, Zhiqiang Lou, Fabian Manhardt et al.CVPR 2022 · 141 citations
- Proposal-Free Video Grounding with Contextual Pyramid NetworkKun Li, Dan Guo, Meng WangAAAI 2021 · 138 citations
- CAPTRA: CAtegory-level Pose Tracking for Rigid and Articulated Objects from Point CloudsYijia Weng, He Wang, Qiang Zhou, Yuzhe Qin et al.ICCV 2021 · 119 citations
- OakInk: A Large-scale Knowledge Repository for Understanding Hand-Object InteractionLixin Yang, Kailin Li, Xinyu Zhan, Fei Wu et al.CVPR 2022 · 79 citations
- AKB-48: A Real-World Articulated Object Knowledge BaseLiu Liu, Wenqiang Xu, Haoyuan Fu, Sucheng Qian et al.CVPR 2022 · 64 citations
Related papers
- VoCAPTER: Voting-based Pose Tracking for Category-level Articulated Object via Inter-frame PriorsLi Zhang, Zean Han, Yan Zhong, Qiaojun Yu et al.ACM MM 2024 · 6 citations
- R^2-Art: Category-Level Articulation Pose Estimation from Single RGB Image via Cascade Render StrategyLi Zhang, Haonan Jiang, Yukang Huo, Yan Zhong et al.AAAI 2025 · 6 citations
- Exploring Category-level Articulated Object Pose Tracking on SE(3) ManifoldsXianhui Meng, Yukang Huo, Li Zhang, Liu Liu et al.AAAI 2026 · 1 citation
- EfficientCAPER: An End-to-End Framework for Fast and Robust Category-Level Articulated Object Pose EstimationXinyi Yu, Haonan Jiang, Li Zhang, Lin Yuanbo Wu et al.NeurIPS 2024
- Instance-Adaptive and Geometric-Aware Keypoint Learning for Category-Level 6D Object Pose EstimationXiao Lin, Wenfei Yang, Yuan Gao, Tianzhu ZhangCVPR 2024
