Neural Assets: 3D-Aware Multi-Object Scene Synthesis with Image Diffusion Models
Ziyi Wu, Yulia Rubanova, Rishabh Kabra, Drew A. Hudson, Igor Gilitschenski, Yusuf Aytar, Sjoerd van Steenkiste, Kelsey R. Allen, Thomas Kipf
摘要
We address the problem of multi-object 3D pose control in image diffusion models. Instead of conditioning on a sequence of text tokens, we propose to use a set of per-object representations, Neural Assets, to control the 3D pose of individual objects in a scene. Neural Assets are obtained by pooling visual representations of objects from a reference image, such as a frame in a video, and are trained to reconstruct the respective objects in a different image, e.g., a later frame in the video. Importantly, we encode object visuals from the reference image while conditioning on object poses from the target frame. This enables learning disentangled appearance and pose features. Combining visual and 3D pose representations in a sequence-of-tokens format allows us to keep the text-to-image architecture of existing models, with Neural Assets in place of text tokens. By fine-tuning a pre-trained text-to-image diffusion model with this information, our approach enables fine-grained 3D pose and placement control of individual objects in a scene. We further demonstrate that Neural Assets can be transferred and recomposed across different scenes. Our model achieves state-of-the-art multi-object editing results on both synthetic 3D scene datasets, as well as two real-world video datasets (Objectron, Waymo Open).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- Orient Anything V2: Unifying Orientation and Rotation UnderstandingZehan Wang, Ziang Zhang, Jiayang Xu, Jialei Wang 等NeurIPS 2025 · 被引用 24 次
- SceneDesigner: Controllable Multi-Object Image Generation with 9-DoF Pose ManipulationZhenyuan Qin, Xincheng Shuai, Henghui DingNeurIPS 2025 · 被引用 11 次
- Orchid: Image Latent Diffusion for Joint Appearance and Geometry GenerationAkshay Krishnan, Xinchen Yan, Vincent Casser, Abhijit KunduICCV 2025 · 被引用 8 次
- Coupled Diffusion Sampling for Training-Free Multi-View Image EditingHadi Alzayer, Yunzhi Zhang, Chen Geng, Jia-Bin Huang 等CVPR 2026 · 被引用 6 次
- InterGSEdit: Interactive 3D Gaussian Splatting Editing with 3D Geometry-Consistent Attention PriorMinghao Wen, Shengjie Wu, Kangkan Wang, Dong LiangICCV 2025 · 被引用 5 次
它引用的顶会 Paper74
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
相关 Paper
- VD3D: Taming Large Video Diffusion Transformers for 3D Camera ControlSherwin Bahmani, Ivan Skorokhodov, Aliaksandr Siarohin, Willi Menapace 等ICLR 2025
- EgoControl: Controllable Egocentric Video Generation via 3D Full-Body PosesEnrico Pallotta, Sina Mokhtarzadeh Azar, Lars Doorenbos, Serdar Ozsoy 等CVPR 2026 · 被引用 7 次
- DriveEditor: A Unified 3D Information-Guided Framework for Controllable Object Editing in Driving ScenesYiyuan Liang, Zhiying Yan, Liqun Chen, Jiahuan Zhou 等AAAI 2025 · 被引用 16 次
- Laconic: A 3D Layout Adapter for Controllable Image CreationLéopold Maillard, Tom Durand, Adrien Ramanana Rahary, Maks OvsjanikovICCV 2025
- DORSal: Diffusion for Object-centric Representations of Scenes et alAllan Jabri, Sjoerd van Steenkiste, Emiel Hoogeboom, Mehdi S. M. Sajjadi 等ICLR 2024 · 被引用 18 次
