DICArt: Advancing Category-level Articulated Object Pose Estimation in Discrete State-Spaces
Li Zhang, Mingyu Mei, Ailing Wang, Xianhui Meng, Yan Zhong, Xinyuan Song, Liu Liu, Rujing Wang, Zaixing He, Cewu Lu
摘要
Articulated object pose estimation is a core task in embodied AI and computer vision. Existing methods typically regress poses in a continuous space, but often struggle with 1) navigating a large, complex search space and 2) failing to incorporate intrinsic kinematic constraints. In this paper, we introduce DICArt (DIsCrete Diffusion for Articulated Object Pose Estimation), a novel framework that formulates pose estimation as a conditional discrete diffusion process. Instead of operating in a continuous domain, DICArt progressively denoises a noisy pose representation through a learned reverse diffusion procedure to recover the ground-truth pose.To improve modeling fidelity, we propose a flexible flow decider that dynamically determines whether each token should be denoised or reset, effectively balancing the real and noise distributions during diffusion. Additionally, we incorporate a hierarchical kinematic coupling strategy, estimating the pose of each rigid part hierarchically to respect the object's kinematic structure.We validate DICArt on both synthetic and real-world datasets with multi-hinged articulated objects. Experimental results demonstrate its superior performance and robustness over state-of-the-art baselines. By integrating discrete generative modeling with structural priors, DICArt offers a new paradigm for reliable category-level 6D pose estimation in complex environments. Codewill be publicly available upon acceptance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper25
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Argmax Flows and Multinomial Diffusion: Learning Categorical DistributionsEmiel Hoogeboom, Didrik Nielsen, Priyank Jaini, Patrick Forré 等NeurIPS 2021 · 被引用 782 次
- AudioLDM: Text-to-Audio Generation with Latent Diffusion ModelsHaohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei 等ICML 2023 · 被引用 773 次
- Vector Quantized Diffusion Model for Text-to-Image SynthesisShuyang Gu, Dong Chen, Jianmin Bao, Fang Wen 等CVPR 2022 · 被引用 607 次
- Listen, Denoise, Action! Audio-Driven Motion Synthesis with Diffusion ModelsSimon Alexanderson, Rajmund Nagy, Jonas Beskow, Gustav Eje HenterSIGGRAPH 2023 · 被引用 191 次
相关 Paper
- R^2-Art: Category-Level Articulation Pose Estimation from Single RGB Image via Cascade Render StrategyLi Zhang, Haonan Jiang, Yukang Huo, Yan Zhong 等AAAI 2025 · 被引用 6 次
- Generative Category-level Object Pose Estimation via Diffusion ModelsJiyao Zhang, Mingdong Wu, Hao DongNeurIPS 2023 · 被引用 65 次
- : Discrete Diffusion Model for Occluded 3D Human Pose EstimationWeiquan Wang, Jun Xiao, Chunping Wang, Wei Liu 等NeurIPS 2024 · 被引用 4 次
- KPA-Tracker: Towards Robust and Real-Time Category-Level Articulated Object 6D Pose TrackingLiu Liu, Anran Huang, Qi Wu, Dan Guo 等AAAI 2024 · 被引用 7 次
- DiffPose: Toward More Reliable 3D Pose EstimationJia Gong, Lin Geng Foo, Zhipeng Fan, Qiuhong Ke 等CVPR 2023
