Two by Two: Learning Multi-Task Pairwise Objects Assembly for Generalizable Robot Manipulation
Yu Qi, Yuanchen Ju, Tianming Wei, Chi Chu, Lawson L. S. Wong, Huazhe Xu
Abstract
3D assembly tasks, such as furniture assembly and component fitting, play a crucial role in daily life and represent essential capabilities for future home robots. Existing benchmarks and datasets predominantly focus on assembling geometric fragments or factory parts, which fall short in addressing the complexities of everyday object interactions and assemblies. To bridge this gap, we present 2BY2, a large-scale annotated dataset for daily pairwise objects assembly, covering 18 fine-grained tasks that reflect real-life scenarios, such as plugging into sockets, arranging flowers in vases, and inserting bread into toasters. 2BY2 dataset includes 1,034 instances and 517 pairwise objects with pose and symmetry annotations, requiring approaches that align geometric shapes while accounting for functional and spatial relationships between objects. Leveraging the 2BY2 dataset, we propose a two-step SE(3) pose estimation method with equivariant features for assembly constraints. Compared to previous shape assembly methods, our approach achieves state-of-the-art performance across all 18 tasks in the 2BY2 dataset. Additionally, robot experiments further validate the reliability and generalization ability of our method for complex 3D assembly tasks. More details and demonstrations can be found at https://tea-lab.github.io/TwoByTwo/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Rectified Point Flow: Generic Point Cloud Pose EstimationTao Sun, Liyuan Zhu, Shengyu Huang, Shuran Song et al.NeurIPS 2025 · 14 citations
- 3D Equivariant Visuomotor Policy Learning via Spherical ProjectionBoce Hu, Dian Wang, David Klee, Heng Tian et al.NeurIPS 2025 · 9 citations
- EquAct: An SE(3)-Equivariant Multi-Task Transformer for 3D Robotic ManipulationXupeng Zhu, Yu Qi, Yizhe Zhu, Robin Walters et al.ICLR 2026 · 9 citations
- MomaGraph: State-Aware Unified Scene Graphs with Vision-Language Models for Embodied Task PlanningYuanchen Ju, Yongyuan Liang, Yen-Jen Wang, Nandiraju Gireesh et al.ICLR 2026 · 5 citations
- EnergyAction: Unimanual to Bimanual Composition with Energy-Based ModelsMingchen Song, Xiang Deng, Jie Wei, Dongmei Jiang et al.CVPR 2026 · 1 citation
Builds on16
- Vector Neurons: A General Framework for SO(3)-Equivariant NetworksCongyue Deng, Or Litany, Yueqi Duan, Adrien Poulenard et al.ICCV 2021 · 411 citations
- Generative 3D Part Assembly via Dynamic Graph LearningGuanqi Zhan, Qingnan Fan, Kaichun Mo, Lin Shao et al.NeurIPS 2020 · 113 citations
- Jigsaw: Learning to Assemble Multiple Fractured ObjectsJiaxin Lu, Yifan Sun, Qixing HuangNeurIPS 2023 · 44 citations
- Neural Shape Mating: Self-Supervised Object Assembly with Adversarial Shape PriorsYun-Chun Chen, Haoda Li, Dylan Turpin, Alec Jacobson et al.CVPR 2022 · 34 citations
- Leveraging SE(3) Equivariance for Learning 3D Geometric Shape AssemblyRuihai Wu, Chenrui Tie, Yushi Du, Yan Zhao et al.ICCV 2023 · 34 citations
Related papers
- AssemblyHands: Towards Egocentric Activity Understanding via 3D Hand Pose EstimationTakehiko Ohkawa, Kun He, Fadime Sener, Tomas Hodan et al.CVPR 2023
- AssemblyBench: Physics-Aware Assembly of Complex Industrial ObjectsDanrui Li, Jiahao Zhang, Bernhard Egger, Moitreya Chatterjee et al.CVPR 2026 · 2 citations
- SceneFun3D: Fine-Grained Functionality and Affordance Understanding in 3D ScenesAlexandros Delitzas, Ayça Takmaz, Federico Tombari, Robert W. Sumner et al.CVPR 2024
- CAD-Estate: Large-scale CAD Model Annotation in RGB VideosKevis-Kokitsi Maninis, Stefan Popov, Matthias Nießner, Vittorio FerrariICCV 2023 · 14 citations
- PhoCaL: A Multi-Modal Dataset for Category-Level Object Pose Estimation with Photometrically Challenging ObjectsPengyuan Wang, HyunJun Jung, Yitong Li, Siyuan Shen et al.CVPR 2022 · 44 citations
