Lune

CVPR2025Top-tier venue

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation

Zehuan Huang, Yuan-Chen Guo, Xingqiao An, Yunhan Yang, Yangguang Li, Zi-Xin Zou, Ding Liang, Xihui Liu, Yan-Pei Cao, Lu Sheng

2025Year
8Top-tier citations

Abstract

Total3D InstPIFu SSR DiffCAD Gen3DSR REPARO Ours (a) Scene generation (b) Generalization Synthetic Data Real-world Image Stylized Image Figure 1. MIDI generates compositional 3D scenes from a single image by extending pre-trained image-to-3D object generation models to multi-instance diffusion models, incorporating a novel multi-instance attention mechanism that captures inter-object interactions. (a) shows our generated scenes compared with those reconstructed by existing methods. (b) presents our generated results on synthetic data, real-world images, and stylized images.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext ca771648-3b21-4e49-9dfd-cecb57d7c79a

Cited by top-tier papers8

Ask how each one uses it

Builds on44

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines