SpatialHand: Generative Object Manipulation from 3D Prespective
Zehan Wang, Jialei Wang, Siyu Chen, Ziang Zhang, Luping Liu, Xize Cheng, Kaihang Pan, Hengshuang Zhao, Zhou Zhao
摘要
We introduce SpatialHand, a novel framework for generative object insertion with precise 3D control. Current generative object manipulation methods primarily operate within the 2D image plane, but often fail to grasp 3D scene complexities, leading to ambiguities in an object's 3D position, orientation, and occlusion relations. SpatialHand addresses this by conceptualizing object insertion from a true ``3D perspective," enabling manipulation with a complete 6 Degrees-of-Freedom (6DoF) controllability. Specifically, our solution naturally and implicitly encodes the 6DoF pose condition by decomposing it into 2D location (via masked image), depth (via composited depth map), and 3D orientation (embedded into latent features). To overcome the scarcity of paired training data, we develop an automated data construction pipeline using synthetic 3D assets, rendering, and subject-driven generation, complemented by visual foundation models for pose estimation. We further design a multi-stage training scheme to progressively drive SpatialHand to robustly follow multiple complex conditions. Extensive experiments reveal our approach's superiority over existing alternatives and its great potential for enabling more versatile and intuitive AR/VR-like object manipulation within images.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
相关 Paper
- HOIDiffusion: Generating Realistic 3D Hand-Object Interaction DataMengqi Zhang, Yang Fu, Zheng Ding, Sifei Liu 等CVPR 2024 · 被引用 18 次
- Hand-Object Interaction Image GenerationHezhen Hu, Weilun Wang, Wengang Zhou, Houqiang LiNeurIPS 2022 · 被引用 24 次
- Direct 3D-Aware Object Insertion via Decomposed Visual ProxiesJingbo Gong, Yikai Wang, Yushi Lan, Yuhao Wan 等ICML 2026 · 被引用 3 次
- SesaHand: Enhancing 3D Hand Reconstruction via Controllable Generation with Semantic and Structural AlignmentZhuoran Zhao, Xianghao Kong, Linlin Yang, Zheng Wei 等ICLR 2026
- SoFar: Language-Grounded Orientation Bridges Spatial Reasoning and Object ManipulationZekun Qi, Wenyao Zhang, Yufei Ding, Runpei Dong 等NeurIPS 2025 · 被引用 65 次
