Predict-Optimize-Distill: A Self-Improving Cycle for 4D Object Understanding
Mingxuan Wu, Huang Huang, Justin Kerr, Chung Min Kim, Anthony Zhang, Brent Yi, Angjoo Kanazawa
摘要
Humans can resort to long-form inspection to build intuition on predicting the 3D configurations of unseen objects. The more we observe the object motion, the better we get at predicting its 3D state immediately. Existing systems either optimize underlying representations from multi-view observations or train a feed-forward predictor from supervised datasets. We introduce Predict-Optimize-Distill (POD), a self-improving framework that interleaves prediction and optimization in a mutually reinforcing cycle to achieve better 4D object understanding with increasing observation time. Given a multi-view object scan and a long-form monocular video of human-object interaction, POD iteratively trains a neural network to predict local part poses from RGB frames, uses this predictor to initialize a global optimization which refines output poses through inverse rendering, then finally distills the results of optimization back into the model by generating synthetic self-labeled training data from novel viewpoints. Each iteration improves both the predictive model and the optimized motion trajectory, creating a virtuous cycle that bootstraps its own training data to learn about the pose configurations of an object. We also introduce a quasi-multiview mining strategy for reducing depth ambiguity by leveraging long video. We evaluate POD on 14 real-world and 5 synthetic objects with various joint types, including revolute and prismatic joints as well as multi-body configurations where parts detach or reattach independently. POD demonstrates significant improvement over a pure optimization baseline which gets stuck in local minima, particularly for longer videos. We also find that POD's performance improves with both video length and successive iterations of the self-improving cycle, highlighting its ability to scale performance with additional observations and looped refinement.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Articulation in Motion: Prior-free Part Mobility Analysis for Articulated Objects By Dynamic-Static DisentanglementHao Ai, Wenjie Chang, Jianbo Jiao, Ales Leonardis 等ICLR 2026 · 被引用 7 次
- ArtHOI: Taming Foundation Models for Monocular 4D Reconstruction of Hand-Articulated-Object InteractionsZikai Wang, Zhilu Zhang, Yiqing Wang, Hui Li 等CVPR 2026 · 被引用 4 次
- FreeArtGS: Articulated Gaussian Splatting Under Free-moving ScenarioHang Dai, Hongwei Fan, Han Zhang, Duojin Wu 等CVPR 2026 · 被引用 3 次
- FUN REC * Reconstructing Functional 3D Scenes from Egocentric Interaction VideosAlexandros Delitzas, Chenyangguang Zhang, Alexey Gavryushin, Tommaso Di Mario 等CVPR 2026
- Kinematic KitbashingMinghao Guo, Victor B. Zordan, Sheldon Andrews, Wojciech Matusik 等SIGGRAPH 2026
它引用的顶会 Paper30
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 被引用 5,687 次
- Nerfies: Deformable Neural Radiance FieldsKeunhong Park, Utkarsh Sinha, Jonathan T. Barron, Sofien Bouaziz 等ICCV 2021 · 被引用 1,442 次
- Learning to Reconstruct 3D Human Pose and Shape via Model-Fitting in the LoopNikos Kolotouros, Georgios Pavlakos, Michael J. Black, Kostas DaniilidisICCV 2019 · 被引用 1,139 次
- STaR: Bootstrapping Reasoning With ReasoningEric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. GoodmanNeurIPS 2022 · 被引用 1,126 次
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu 等CVPR 2024 · 被引用 847 次
相关 Paper
- CARI4D: Category Agnostic 4D Reconstruction of Human-Object InteractionXianghui Xie, Bowen Wen, Yan Chang, Hesam Rabeti 等CVPR 2026 · 被引用 16 次
- Inferring Compositional 4D Scenes without Ever Seeing OneAhmet Berke Gökmen, Ajad Chhatkuli, Luc Van Gool, Danda PaudelCVPR 2026 · 被引用 1 次
- DreamScene4D: Dynamic Multi-Object Scene Generation from Monocular VideosWen-Hsuan Chu, Lei Ke, Katerina FragkiadakiNeurIPS 2024 · 被引用 75 次
- UFO: Unifying Feed-Forward and Optimization-based Methods for Large Driving Scene ModelingKaiyuan Tan, Yingying Shen, Ziyue Zhu, Mingfei Tu 等CVPR 2026 · 被引用 4 次
- Revealing Occlusions with 4D Neural FieldsBasile Van Hoorick, Purva Tendulkar, Dídac Surís, Dennis Park 等CVPR 2022 · 被引用 10 次
