Articulate-Anything: Automatic Modeling of Articulated Objects via a Vision-Language Foundation Model
Long Le, Jason Xie, William Liang, Hung-Ju Wang, Yue Yang, Yecheng Jason Ma, Kyle Vedder, Arjun Krishna, Dinesh Jayaraman, Eric Eaton
Abstract
Interactive 3D simulated objects are crucial in AR/VR, animations, and robotics, driving immersive experiences and advanced automation. However, creating these articulated objects requires extensive human effort and expertise, limiting their broader applications. To overcome this challenge, we present ARTICULATE-ANYTHING, a system that automates the articulation of diverse, complex objects from many input modalities, including text, images, and videos. ARTICULATE-ANYTHING leverages vision-language models (VLMs) to generate code that can be compiled into an interactable digital twin for use in standard 3D simulators. Our system exploits existing 3D asset datasets via a mesh retrieval mechanism, along with an actor-critic system that iteratively proposes, evaluates, and refines solutions for articulating the objects, self-correcting errors to achieve a robust outcome. Qualitative evaluations demonstrate ARTICULATE-ANYTHING's capability to articulate complex and even ambiguous object affordances by leveraging rich grounded inputs. In extensive quantitative experiments on the standard PartNet-Mobility dataset, ARTICULATE-ANYTHING substantially outperforms prior work, increasing the success rate from 8.7-12.2% to 75% and setting a new bar for state-of-the-art performance. We further showcase the utility of our system by generating 3D assets from in-the-wild video inputs, which are then used to train robotic policies for fine-grained manipulation tasks in simulation that go beyond basic pick and place. These policies are then transferred to a real robotic system. Published as a conference paper at ICLR 2025 simulators that scale to millions of FPS and hundreds of GPUs (Xiang et al., 2020; Makoviychuk et al., 2021) , enabling policy learning on a staggering scale. However, a critical bottleneck in this research direction persists: the immense human labor required to construct realistic, interactable environments for these agents to learn within. Despite the existence of large, open libraries of static object geometries -with the largest open dataset containing over 10 million objects (Deitke et al., 2024) -we have comparatively minuscule open libraries of articulated 3D objects (only around 2,300 objects (Xiang et al., 2020) ). This scarcity stems from the time-consuming, labor-intensive, and expertise-demanding nature of the manual annotation process. To address this challenge, we present ARTICULATE-ANYTHING, a novel approach in automatic articulation that harnesses the power of leading foundation vision-language models (VLMs) to articulate a diverse range of objects of arbitrary complexity through iterative feedback (Fig. 1 ). ARTICULATE-ANYTHING represents a step function improvement in quality, accuracy (8.7-12.2% to 75%), and generalizability over prior art (Chen et al., 2024; Mandi et al., 2024) , overcoming previous limitations that restricted success to only a narrow range of object categories and joint types. Unlike prior art, which has been limited by the impoverished input of bounding boxes or static images, ARTICULATE-ANYTHING affords the flexibility of consuming rich, grounded inputs from text, images, or even videos, enabling users to request exotic articulation descriptions or resolve articulation ambiguities. For example, the right column of Fig. 7 features a digital model of a window that could plausibly slide or tip to open; when ARTICULATE-ANYTHING is shown an in-the-wild video demonstration, it accurately produces the desired sliding motion. To achieve this level of flexibility and accuracy, ARTICULATE-ANYTHING employs an actor-critic system with two core components: (1) a vision-language actor that synthesizes high-level Python code, which can be compiled into Unified Robot Description Format (URDF) files and (2) a vision-language critic that provides feedback on the rendered prediction compared against available ground-truth. The result is an agentic system that can automatically self-evaluate and iteratively improve the articulation of complex objects. Beyond robotics, the flexibility of ARTICULATE-ANYTHING's inputs married with its high-quality outputs puts automatic generation of rich, high-quality, and diverse virtual environments within reach with broad-reaching applications to 3D/VR (Kim et al., 2024 ), human-computer interaction (Jiang et al., 2023), and animation (Yang et al., 2022).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2091bc0f-40cf-4200-85b0-6fdb82eaa1d3Cited by top-tier papers25
- PhysX-3D: Physical-Grounded 3D Asset GenerationZiang Cao, Zhaoxi Chen, Liang Pan, Ziwei LiuNeurIPS 2025 · 46 citations
- PhysX-Anything: Simulation-Ready Physical 3D Assets from Single ImageZiang Cao, Fangzhou Hong, Zhaoxi Chen, Liang Pan et al.CVPR 2026 · 36 citations
- URDF-Anything: Constructing Articulated Objects with 3D Multimodal Language ModelZhe Li, Xiang Bai, Jieyu Zhang, Zhuangzhe Wu et al.NeurIPS 2025 · 24 citations
- Particulate: Feed-Forward 3D Object ArticulationRuining Li, Yuxin Yao, Chuanxia Zheng, Christian Rupprecht et al.CVPR 2026 · 22 citations
- ArtVIP: Articulated Digital Assets of Visual Realism, Modular Interaction, and Physical Fidelity for Robot LearningZhao Jin, Zhengping Che, Tao Li, Zhen Zhao et al.ICLR 2026 · 15 citations
Builds on11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- PARIS: Part-level Reconstruction and Motion Analysis for Articulated ObjectsJiayi Liu, Ali Mahdavi-Amiri, Manolis SavvaICCV 2023 · 103 citations
- Full-Body Articulated Human-Object InteractionNan Jiang, Tengyu Liu, Zhexuan Cao, Jieming Cui et al.ICCV 2023 · 80 citations
- Ditto: Building Digital Twins of Articulated Objects from InteractionZhenyu Jiang, Cheng-Chun Hsu, Yuke ZhuCVPR 2022 · 77 citations
- NAP: Neural 3D Articulated Object PriorJiahui Lei, Congyue Deng, William B. Shen, Leonidas J. Guibas et al.NeurIPS 2023 · 53 citations
Related papers
- ArtLLM: Generating Articulated Assets via 3D LLMPenghao Wang, Siyuan Xie, Hongyu Yan, Xianghui Yang et al.CVPR 2026 · 7 citations
- Real2Code: Reconstruct Articulated Objects via Code GenerationZhao Mandi, Yijia Weng, Dominik Bauer, Shuran SongICLR 2025
- Arti-PG: A Toolbox for Procedurally Synthesizing Large-Scale and Diverse Articulated Objects with Rich AnnotationsJianhua Sun, Yuxuan Li, Jiude Wei, Longfei Xu et al.ICCV 2025 · 3 citations
- Articulate3D: Holistic Understanding of 3D Scenes as Universal Scene DescriptionAnna-Maria Halacheva, Yang Miao, Jan-Nico Zaech, Xi Wang et al.ICCV 2025 · 2 citations
- SPARK: Sim-ready Part-level Articulated Reconstruction with VLM KnowledgeYumeng He, Ying Jiang, Jiayin Lu, Yin Yang et al.CVPR 2026 · 6 citations
