Two by Two: Learning Multi-Task Pairwise Objects Assembly for Generalizable Robot Manipulation
Yu Qi, Yuanchen Ju, Tianming Wei, Chi Chu, Lawson L. S. Wong, Huazhe Xu
摘要
3D assembly tasks, such as furniture assembly and component fitting, play a crucial role in daily life and represent essential capabilities for future home robots. Existing benchmarks and datasets predominantly focus on assembling geometric fragments or factory parts, which fall short in addressing the complexities of everyday object interactions and assemblies. To bridge this gap, we present 2BY2, a large-scale annotated dataset for daily pairwise objects assembly, covering 18 fine-grained tasks that reflect real-life scenarios, such as plugging into sockets, arranging flowers in vases, and inserting bread into toasters. 2BY2 dataset includes 1,034 instances and 517 pairwise objects with pose and symmetry annotations, requiring approaches that align geometric shapes while accounting for functional and spatial relationships between objects. Leveraging the 2BY2 dataset, we propose a two-step SE(3) pose estimation method with equivariant features for assembly constraints. Compared to previous shape assembly methods, our approach achieves state-of-the-art performance across all 18 tasks in the 2BY2 dataset. Additionally, robot experiments further validate the reliability and generalization ability of our method for complex 3D assembly tasks. More details and demonstrations can be found at https://tea-lab.github.io/TwoByTwo/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Rectified Point Flow: Generic Point Cloud Pose EstimationTao Sun, Liyuan Zhu, Shengyu Huang, Shuran Song 等NeurIPS 2025 · 被引用 14 次
- 3D Equivariant Visuomotor Policy Learning via Spherical ProjectionBoce Hu, Dian Wang, David Klee, Heng Tian 等NeurIPS 2025 · 被引用 9 次
- EquAct: An SE(3)-Equivariant Multi-Task Transformer for 3D Robotic ManipulationXupeng Zhu, Yu Qi, Yizhe Zhu, Robin Walters 等ICLR 2026 · 被引用 9 次
- MomaGraph: State-Aware Unified Scene Graphs with Vision-Language Models for Embodied Task PlanningYuanchen Ju, Yongyuan Liang, Yen-Jen Wang, Nandiraju Gireesh 等ICLR 2026 · 被引用 5 次
- EnergyAction: Unimanual to Bimanual Composition with Energy-Based ModelsMingchen Song, Xiang Deng, Jie Wei, Dongmei Jiang 等CVPR 2026 · 被引用 1 次
它引用的顶会 Paper16
- Vector Neurons: A General Framework for SO(3)-Equivariant NetworksCongyue Deng, Or Litany, Yueqi Duan, Adrien Poulenard 等ICCV 2021 · 被引用 411 次
- Generative 3D Part Assembly via Dynamic Graph LearningGuanqi Zhan, Qingnan Fan, Kaichun Mo, Lin Shao 等NeurIPS 2020 · 被引用 113 次
- Jigsaw: Learning to Assemble Multiple Fractured ObjectsJiaxin Lu, Yifan Sun, Qixing HuangNeurIPS 2023 · 被引用 44 次
- Neural Shape Mating: Self-Supervised Object Assembly with Adversarial Shape PriorsYun-Chun Chen, Haoda Li, Dylan Turpin, Alec Jacobson 等CVPR 2022 · 被引用 34 次
- Leveraging SE(3) Equivariance for Learning 3D Geometric Shape AssemblyRuihai Wu, Chenrui Tie, Yushi Du, Yan Zhao 等ICCV 2023 · 被引用 34 次
相关 Paper
- AssemblyHands: Towards Egocentric Activity Understanding via 3D Hand Pose EstimationTakehiko Ohkawa, Kun He, Fadime Sener, Tomas Hodan 等CVPR 2023
- AssemblyBench: Physics-Aware Assembly of Complex Industrial ObjectsDanrui Li, Jiahao Zhang, Bernhard Egger, Moitreya Chatterjee 等CVPR 2026 · 被引用 2 次
- SceneFun3D: Fine-Grained Functionality and Affordance Understanding in 3D ScenesAlexandros Delitzas, Ayça Takmaz, Federico Tombari, Robert W. Sumner 等CVPR 2024
- CAD-Estate: Large-scale CAD Model Annotation in RGB VideosKevis-Kokitsi Maninis, Stefan Popov, Matthias Nießner, Vittorio FerrariICCV 2023 · 被引用 14 次
- PhoCaL: A Multi-Modal Dataset for Category-Level Object Pose Estimation with Photometrically Challenging ObjectsPengyuan Wang, HyunJun Jung, Yitong Li, Siyuan Shen 等CVPR 2022 · 被引用 44 次
