Neural Assembler: Learning to Generate Fine-Grained Robotic Assembly Instructions from Multi-View Images
Hongyu Yan, Yadong Mu
Abstract
Image-guided object assembly represents a burgeoning research topic in computer vision. This paper introduces a novel task: translating multi-view images of a structural 3D model (for example, one constructed with building blocks drawn from a 3D-object library) into a detailed sequence of assembly instructions executable by a robotic arm. Fed with multi-view images of the target 3D model for replication, the model designed for this task must address several sub-tasks, including recognizing individual components used in constructing the 3D model, estimating the geometric pose of each component, and deducing a feasible assembly order adhering to physical rules. Establishing accurate 2D-3D correspondence between multi-view images and 3D objects is technically challenging. To tackle this, we propose an end-to-end model known as the Neural Assembler. This model learns an object graph where each vertex represents recognized components from the images, and the edges specify the topology of the 3D model, enabling the derivation of an assembly plan. We establish benchmarks for this task and conduct comprehensive empirical evaluations of Neural Assembler and alternative solutions. Our experiments clearly demonstrate the superiority of Neural Assembler.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f17ca079-2b66-4bf2-b101-05e4d02c9aceBuilds on10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Image Segmentation Using Text and Image PromptsTimo Lüddecke, Alexander S. EckerCVPR 2022 · 457 citations
- Generative 3D Part Assembly via Dynamic Graph LearningGuanqi Zhan, Qingnan Fan, Kaichun Mo, Lin Shao et al.NeurIPS 2020 · 113 citations
- SGTR: End-to-end Scene Graph Generation with TransformerRongjie Li, Songyang Zhang, Xuming HeCVPR 2022 · 108 citations
- FS-DETR: Few-Shot DEtection TRansformer with prompting and without re-trainingAdrian Bulat, Ricardo Guerrero, Brais Martínez, Georgios TzimiropoulosICCV 2023 · 61 citations
Related papers
- Imagine: Image-Guided 3D Part Assembly with Structure Knowledge GraphWeihao Wang, Yu Lan, Mingyu You, Bin HeAAAI 2025
- Category-Level Multi-Part Multi-Joint 3D Shape AssemblyYichen Li, Kaichun Mo, Yueqi Duan, He Wang et al.CVPR 2024
- Manual-PA: Learning 3D Part Assembly from Instruction DiagramsJiahao Zhang, Anoop Cherian, Cristian Rodriguez, Weijian Deng et al.ICCV 2025 · 2 citations
- Completing 3D Partial Assemblies with View-Consistent 2D-3D CorrespondenceWeihao Wang, Yu Lan, Mingyu You, Bin HeICCV 2025 · 1 citation
- DiffAssemble: A Unified Graph-Diffusion Model for 2D and 3D ReassemblyGianluca Scarpellini, Stefano Fiorini, Francesco Giuliari, Pietro Morerio et al.CVPR 2024 · 12 citations
