AssemblyBench: Physics-Aware Assembly of Complex Industrial Objects
Danrui Li, Jiahao Zhang, Bernhard Egger, Moitreya Chatterjee, Suhas Lohit, Tim K. Marks, Anoop Cherian
Abstract
Introduction Assembling objects from parts requires understanding multimodal instructions, linking them to 3D components, and predicting physically plausible 6-DoF motions for each assembly step. Existing datasets focus on simplified scenarios, overlooking shape complexities and assembly trajectories in industrial assemblies. We introduce AssemblyBench, a synthetic dataset of 2,789 industrial objects with multimodal instruction manuals, corresponding 3D part models, and part assembly trajectories. AssemblyBench is designed to facilitate research that bridges instructional manual understanding and the execution of assembly steps, serving as a benchmark for the development of next-generation assembly algorithms. All necessary data for training, validation, and testing are publicly released. At a Glance - Contents: - Total assemblies: 2,789 - Dataset file size: 2.3 GB (zipped), 4.6 GB (extracted) - Split sizes: - Train: 2,231 (all.train.txt) - Val: 278 (all.val.txt) - Test: 280 (all.test.txt) - Steps (parts) per assembly: - Min: 2 - Max: 20 - Mean: 6.7 Other Resources The code associated with the approach will be released separately on GitHub. Citation If you use the AssemblyBench dataset in your research, please cite our contribution: @inproceedingsLi2026AssemblyBench, author = Li, Danrui and Zhang, Jiahao and Egger, Bernhard and Chatterjee, Moitreya and Lohit, Suhas and Marks, Tim K. and Cherian, Anoop, title = AssemblyBench: Physics-Aware Assembly of Complex Industrial Objects, booktitle = IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), year = 2026, License AssemblyBench dataset extends the Assemble-Them-All dataset, originally released under the MIT License. The original data remains under the MIT License. As the Assemble-Them-All dataset uses assets from the Fusion 360 Gallary Dataset, please refer to the Fusion 360 Gallery Dataset License for legal usage. All new annotations and modifications introduced in our release are licensed under the CC-BY-SA-4.0. SPDX-License-Identifier: CC-BY-SA-4.0 Created by Mitsubishi Electric Research Laboratories (MERL), 2026
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on10
- Assembly101: A Large-Scale Multi-View Video Dataset for Understanding Procedural ActivitiesFadime Sener, Dibyadip Chatterjee, Daniel Shelepov, Kun He et al.CVPR 2022 · 168 citations
- Text2CAD: Generating Sequential CAD Designs from Beginner-to-Expert Level Text PromptsMohammad Sadil Khan, Sankalp Sinha, Talha Uddin Sheikh, Didier Stricker et al.NeurIPS 2024 · 148 citations
- CompoNet: Learning to Generate the Unseen by Part Synthesis and CompositionNadav Schor, Oren Katzir, Hao Zhang, Daniel Cohen-OrICCV 2019 · 63 citations
- Generating Physically Stable and Buildable Brick Structures from TextAva Pun, Kangle Deng, Ruixuan Liu, Deva Ramanan et al.ICCV 2025 · 9 citations
- Manual-PA: Learning 3D Part Assembly from Instruction DiagramsJiahao Zhang, Anoop Cherian, Cristian Rodriguez, Weijian Deng et al.ICCV 2025 · 2 citations
Related papers
- Two by Two: Learning Multi-Task Pairwise Objects Assembly for Generalizable Robot ManipulationYu Qi, Yuanchen Ju, Tianming Wei, Chi Chu et al.CVPR 2025
- AssemblyHands: Towards Egocentric Activity Understanding via 3D Hand Pose EstimationTakehiko Ohkawa, Kun He, Fadime Sener, Tomas Hodan et al.CVPR 2023
- Rethinking Video Generation Model for the Embodied WorldYufan Deng, Zilin Pan, Hongyu Zhang, Xiaojie Li et al.ICML 2026 · 24 citations
- R3-Bench: Reproducible Real-world Reverse Engineering Dataset for Symbol RecoveryMuzhi Yu, Zhengran Zeng, Wei Ye, Jinan Sun et al.ASE 2025
- SiM3D: Single-Instance Multiview Multimodal and Multisetup 3D Anomaly Detection BenchmarkAlex Costanzino, Pierluigi Zama Ramirez, Luigi Lella, Matteo Ragaglia et al.ICCV 2025 · 2 citations
