Neural Assembler: Learning to Generate Fine-Grained Robotic Assembly Instructions from Multi-View Images
Hongyu Yan, Yadong Mu
摘要
Image-guided object assembly represents a burgeoning research topic in computer vision. This paper introduces a novel task: translating multi-view images of a structural 3D model (for example, one constructed with building blocks drawn from a 3D-object library) into a detailed sequence of assembly instructions executable by a robotic arm. Fed with multi-view images of the target 3D model for replication, the model designed for this task must address several sub-tasks, including recognizing individual components used in constructing the 3D model, estimating the geometric pose of each component, and deducing a feasible assembly order adhering to physical rules. Establishing accurate 2D-3D correspondence between multi-view images and 3D objects is technically challenging. To tackle this, we propose an end-to-end model known as the Neural Assembler. This model learns an object graph where each vertex represents recognized components from the images, and the edges specify the topology of the 3D model, enabling the derivation of an assembly plan. We establish benchmarks for this task and conduct comprehensive empirical evaluations of Neural Assembler and alternative solutions. Our experiments clearly demonstrate the superiority of Neural Assembler.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Image Segmentation Using Text and Image PromptsTimo Lüddecke, Alexander S. EckerCVPR 2022 · 被引用 457 次
- Generative 3D Part Assembly via Dynamic Graph LearningGuanqi Zhan, Qingnan Fan, Kaichun Mo, Lin Shao 等NeurIPS 2020 · 被引用 113 次
- SGTR: End-to-end Scene Graph Generation with TransformerRongjie Li, Songyang Zhang, Xuming HeCVPR 2022 · 被引用 108 次
- FS-DETR: Few-Shot DEtection TRansformer with prompting and without re-trainingAdrian Bulat, Ricardo Guerrero, Brais Martínez, Georgios TzimiropoulosICCV 2023 · 被引用 61 次
相关 Paper
- Imagine: Image-Guided 3D Part Assembly with Structure Knowledge GraphWeihao Wang, Yu Lan, Mingyu You, Bin HeAAAI 2025
- Category-Level Multi-Part Multi-Joint 3D Shape AssemblyYichen Li, Kaichun Mo, Yueqi Duan, He Wang 等CVPR 2024
- Manual-PA: Learning 3D Part Assembly from Instruction DiagramsJiahao Zhang, Anoop Cherian, Cristian Rodriguez, Weijian Deng 等ICCV 2025 · 被引用 2 次
- Completing 3D Partial Assemblies with View-Consistent 2D-3D CorrespondenceWeihao Wang, Yu Lan, Mingyu You, Bin HeICCV 2025 · 被引用 1 次
- DiffAssemble: A Unified Graph-Diffusion Model for 2D and 3D ReassemblyGianluca Scarpellini, Stefano Fiorini, Francesco Giuliari, Pietro Morerio 等CVPR 2024 · 被引用 12 次
