3D Implicit Transporter for Temporally Consistent Keypoint Discovery
Chengliang Zhong, Yuhang Zheng, Yupeng Zheng, Hao Zhao, Li Yi, Xiaodong Mu, Ling Wang, Pengfei Li, Guyue Zhou, Chao Yang, Xinliang Zhang, Jian Zhao
Abstract
Keypoint-based representation has proven advantageous in various visual and robotic tasks. However, the existing 2D and 3D methods for detecting keypoints mainly rely on geometric consistency to achieve spatial alignment, neglecting temporal consistency. To address this issue, the Transporter method was introduced for 2D data, which reconstructs the target frame from the source frame to incorporate both spatial and temporal information. However, the direct application of the Transporter to 3D point clouds is infeasible due to their structural differences from 2D images. Thus, we propose the first 3D version of the Transporter, which leverages hybrid 3D representation, cross attention, and implicit reconstruction. We apply this new learning system on 3D articulated objects and non-rigid animals (humans and rodents) and show that learned keypoints are spatio-temporally consistent. Additionally, we propose a closed-loop control strategy that utilizes the learned keypoints for 3D object manipulation and demonstrate its superior performance. Codes are available at https://github.com/zhongcl-thu/3D-Implicit-Transporter.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 38223be0-fce9-4e77-bb1b-66718bb14860Cited by top-tier papers8
- SoFar: Language-Grounded Orientation Bridges Spatial Reasoning and Object ManipulationZekun Qi, Wenyao Zhang, Yufei Ding, Runpei Dong et al.NeurIPS 2025 · 65 citations
- RoboCodeX: Multimodal Code Generation for Robotic Behavior SynthesisYao Mu, Junting Chen, Qinglong Zhang, Shoufa Chen et al.ICML 2024 · 50 citations
- FastMAC: Stochastic Spectral Sampling of Correspondence GraphYifei Zhang, Hao Zhao, Hongyang Li, Siheng ChenCVPR 2024 · 18 citations
- Dual-frame Fluid Motion Estimation with Test-time Optimization and Zero-divergence LossYifei Zhang, Huan-ang Gao, Zhou Jiang, Hao ZhaoNeurIPS 2024 · 4 citations
- PartRM: Modeling Part-Level Dynamics with Large Cross-State Reconstruction ModelMingju Gao, Yike Pan, Huan-ang Gao, Zongzheng Zhang et al.CVPR 2025
Builds on21
- Implicit Neural Representations with Periodic Activation FunctionsVincent Sitzmann, Julien N. P. Martel, Alexander W. Bergman, David B. Lindell et al.NeurIPS 2020 · 4,008 citations
- Neural Sparse Voxel FieldsLingjie Liu, Jiatao Gu, Kyaw Zaw Lin, Tat-Seng Chua et al.NeurIPS 2020 · 1,535 citations
- GRAF: Generative Radiance Fields for 3D-Aware Image SynthesisKatja Schwarz, Yiyi Liao, Michael Niemeyer, Andreas GeigerNeurIPS 2020 · 1,001 citations
- Where2Act: From Pixels to Actions for Articulated 3D ObjectsKaichun Mo, Leonidas J. Guibas, Mustafa Mukadam, Abhinav Gupta et al.ICCV 2021 · 240 citations
- USIP: Unsupervised Stable Interest Point Detection From 3D Point CloudsJiaxin Li, Gim Hee LeeICCV 2019 · 206 citations
Related papers
- Unsupervised Learning of Visual 3D Keypoints for ControlBoyuan Chen, Pieter Abbeel, Deepak PathakICML 2021 · 46 citations
- KeyTr: Keypoint Transporter for 3D Reconstruction of Deformable Objects in VideosDavid Novotný, Ignacio Rocco, Samarth Sinha, Alexandre Carlier et al.CVPR 2022 · 11 citations
- KPA-Tracker: Towards Robust and Real-Time Category-Level Articulated Object 6D Pose TrackingLiu Liu, Anran Huang, Qi Wu, Dan Guo et al.AAAI 2024 · 7 citations
- Key-Grid: Unsupervised 3D Keypoints Detection using Grid Heatmap FeaturesChengkai Hou, Zhengrong Xue, Bingyang Zhou, Jinghan Ke et al.NeurIPS 2024 · 9 citations
- 3DVG-Transformer: Relation Modeling for Visual Grounding on Point CloudsLichen Zhao, Daigang Cai, Lu Sheng, Dong XuICCV 2021 · 234 citations
