Monomobility: Zero-Shot 3D Mobility Analysis From Monocular Videos
Hongyi Zhou, Yulan Guo, Xiaogang Wang, Kai Xu
Abstract
Accurately analyzing the motion parts and their motion attributes in dynamic environments is crucial for advancing key areas such as embodied intelligence. Addressing the limitations of existing methods that rely on dense multiview images or detailed part-level annotations, we propose an innovative framework that can analyze 3D mobility from monocular videos in a zero-shot manner. This framework can precisely parse motion parts and motion attributes only using a monocular video, completely eliminating the need for annotated training data. Specifically, our method first constructs the scene geometry and roughly analyzes the motion parts and their initial motion attributes combining depth estimation, optical flow analysis and point cloud registration method, then employs 2D Gaussian splatting for scene representation. Building on this, we introduce an end-to-end dynamic scene optimization algorithm specifically designed for articulated objects, refining the initial analysis results to ensure the system can handle 'rotation', 'translation', and even complex movements ('rota-tion+translation'), demonstrating high flexibility and versatility. To validate the robustness and wide applicability of our method, we created a comprehensive dataset comprising both simulated and real-world scenarios. Experimental results show that our framework can effectively analyze articulated object motions in an annotation-free manner, showcasing its significant potential in future embodied intelligence applications. The project page is at: https: //monomobility.github.io/MonoMobility.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f1d050f9-a793-48ea-82e6-efc80e17e60cCited by top-tier papers5
- FreeArtGS: Articulated Gaussian Splatting Under Free-moving ScenarioHang Dai, Hongwei Fan, Han Zhang, Duojin Wu et al.CVPR 2026 · 3 citations
- Self-Supervised Learning of Hybrid Part-Aware 3D Representations of 2D Gaussians and SuperquadricsZhirui Gao, Renjiao Yi, Yuhang Huang, Wei Chen et al.ICCV 2025 · 2 citations
- Curve-Aware Gaussian Splatting for 3D Parametric Curve ReconstructionZhirui Gao, Renjiao Yi, Yaqiao Dai, Xuening Zhu et al.ICCV 2025 · 1 citation
- Clay-to-Stone: Phase-wise 3D Gaussian Splatting for Monocular Articulated Hand-Object Manipulation ModelingXingyu Liu, Pengfei Ren, Qi Qi, Haifeng Sun et al.CVPR 2026 · 1 citation
- D-Prism: Differentiable Primitives for Structured Dynamic ModelingXingyuan Yu, Yijin Li, Chong Zeng, Yuhang Ming et al.CVPR 2026
Builds on21
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao et al.NeurIPS 2024 · 2,305 citations
- 2D Gaussian Splatting for Geometrically Accurate Radiance FieldsBinbin Huang, Zehao Yu, Anpei Chen, Andreas Geiger et al.SIGGRAPH 2024 · 660 citations
- 4D Gaussian Splatting for Real-Time Dynamic Scene RenderingGuanjun Wu, Taoran Yi, Jiemin Fang, Lingxi Xie et al.CVPR 2024 · 513 citations
- DUSt3R: Geometric 3D Vision Made EasyShuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii et al.CVPR 2024 · 302 citations
Related papers
- Ego3DT: Tracking Every 3D Object in Ego-centric VideosShengyu Hao, Wenhao Chai, Zhonghan Zhao, Meiqi Sun et al.ACM MM 2024 · 6 citations
- MoCaNet: Motion Retargeting In-the-Wild via Canonicalization NetworksWentao Zhu, Zhuoqian Yang, Ziang Di, Wayne Wu et al.AAAI 2022 · 24 citations
- Object-centric 3D Motion Field for Robot Learning from Human VideosZhao-Heng Yin, Sherry Yang, Pieter AbbeelNeurIPS 2025 · 18 citations
- FreeGaussian: Annotation-free Control of Articulated Objects via 3D Gaussian Splats with Flow DerivativesQizhi Chen, Delin Qu, Junli Liu, Yiwen Tang et al.AAAI 2026 · 1 citation
- TRACE: Learning 3D Gaussian Physical Dynamics from Multi-View VideosJinxi Li, Ziyang Song, Bo YangICCV 2025 · 3 citations
