PhysX-Anything: Simulation-Ready Physical 3D Assets from Single Image
Ziang Cao, Fangzhou Hong, Zhaoxi Chen, Liang Pan, Ziwei Liu
Abstract
3D modeling is shifting from static visual representations toward physical, articulated assets that can be directly used in simulation and interaction. However, most existing 3D generation methods overlook key physical and articulation properties, thereby limiting their utility in embodied AI. To bridge this gap, we introduce PhysX-Anything, the first simulation-ready physical 3D generative framework that, given a single in-the-wild image, produces high-quality sim-ready 3D assets with explicit geometry, articulation, and physical attributes. Specifically, we propose the first VLM-based physical 3D generative model, along with a new 3D representation that efficiently tokenizes geometry. It reduces the number of tokens by 193, enabling explicit geometry learning within standard VLM token budgets without introducing any special tokens during fine-tuning and significantly improving generative quality. In addition, to overcome the limited diversity of existing physical 3D datasets, we construct a new dataset, PhysX-Mobility, which expands the object categories in prior physical 3D datasets by over 2 and includes more than 2K common real-world objects with rich physical annotations. Extensive experiments on PhysX-Mobility and in-the-wild images demonstrate that PhysX-Anything delivers strong generative performance and robust generalization. Furthermore, simulation-based experiments in a MuJoCo-style environment validate that our sim-ready assets can be directly used for contact-rich robotic policy learning. We believe PhysX-Anything can substantially empower a broad range of downstream applications, especially in embodied AI and physics-based simulation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 496bc8df-aceb-4716-a017-e932829337a5Cited by top-tier papers4
- PhysX-3D: Physical-Grounded 3D Asset GenerationZiang Cao, Zhaoxi Chen, Liang Pan, Ziwei LiuNeurIPS 2025 · 46 citations
- PhysForge: Generating Physics-Grounded 3D Assets for Interactive Virtual WorldYunhan Yang, Chunshi Wang, Junliang Ye, YANG LI et al.ICML 2026 · 6 citations
- RealTwin: Concept Graph Representation and Grounding Framework for Reality-Preserving Digital Twin ReconstructionZisu Li, Ruohao Li, Jiawei Li, Chao Liu et al.CHI 2026 · 1 citation
- SimArt: Decomposing Monolithic Meshes into Sim-ready Articulated Assets via MLLMChuanrui Zhang, Minghan Qin, Yuang Wang, Baifeng Xie et al.SIGGRAPH 2026
Builds on17
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Efficient Geometry-aware 3D Generative Adversarial NetworksEric R. Chan, Connor Z. Lin, Matthew A. Chan, Koki Nagano et al.CVPR 2022 · 984 citations
- DreamFusion: Text-to-3D using 2D DiffusionBen Poole, Ajay Jain, Jonathan T. Barron, Ben MildenhallICLR 2023 · 463 citations
- MeshGPT: Generating Triangle Meshes with Decoder-Only TransformersYawar Siddiqui, Antonio Alliegro, Alexey Artemov, Tatiana Tommasi et al.CVPR 2024 · 74 citations
- Large-Vocabulary 3D Diffusion Model with TransformerZiang Cao, Fangzhou Hong, Tong Wu, Liang Pan et al.ICLR 2024 · 54 citations
Related papers
- SPARK: Sim-ready Part-level Articulated Reconstruction with VLM KnowledgeYumeng He, Ying Jiang, Jiayin Lu, Yin Yang et al.CVPR 2026 · 6 citations
- Articulate-Anything: Automatic Modeling of Articulated Objects via a Vision-Language Foundation ModelLong Le, Jason Xie, William Liang, Hung-Ju Wang et al.ICLR 2025
- URDF-Anything: Constructing Articulated Objects with 3D Multimodal Language ModelZhe Li, Xiang Bai, Jieyu Zhang, Zhuangzhe Wu et al.NeurIPS 2025 · 24 citations
- DexGrasp Anything: Towards Universal Robotic Dexterous Grasping with Physics AwarenessYiming Zhong, Qi Jiang, Jingyi Yu, Yuexin MaCVPR 2025
- ArtLLM: Generating Articulated Assets via 3D LLMPenghao Wang, Siyuan Xie, Hongyu Yan, Xianghui Yang et al.CVPR 2026 · 7 citations
