Multi-Plane Program Induction with 3D Box Priors
Yikai Li, Jiayuan Mao, Xiuming Zhang, Bill Freeman, Josh Tenenbaum, Noah Snavely, Jiajun Wu
Abstract
We consider two important aspects in understanding and editing images: modeling regular, program-like texture or patterns in 2D planes, and 3D posing of these planes in the scene. Unlike prior work on image-based program synthesis, which assumes the image contains a single visible 2D plane, we present Box Program Induction (BPI), which infers a program-like scene representation that simultaneously models repeated structure on multiple 2D planes, the 3D position and orientation of the planes, and camera parameters, all from a single image. Our model assumes a box prior, i.e., that the image captures either an inner view or an outer view of a box in 3D. It uses neural networks to infer visual cues such as vanishing points or wireframe lines to guide a search-based algorithm to find the program that best explains the image. Such a holistic, structured scene representation enables 3D-aware interactive image editing operations such as inpainting missing pixels, changing camera parameters, and extrapolate the image contents. *: indicates equal contribution. 34th Conference on Neural Information Processing Systems (NeurIPS 2020),
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Worldsheet: Wrapping the World in a 3D Sheet for View Synthesis from a Single ImageRonghang Hu, Nikhila Ravi, Alexander C. Berg, Deepak PathakICCV 2021 · 97 citations
- PlaneTR: Structure-Guided Transformers for 3D Plane RecoveryBin Tan, Nan Xue, Song Bai, Tianfu Wu et al.ICCV 2021 · 51 citations
- Look Outside the Room: Synthesizing A Consistent Long-Term 3D Scene Video from A Single ImageXuanchi Ren, Xiaolong WangCVPR 2022 · 42 citations
- Structure from Duplicates: Neural Inverse Graphics from a Pile of ObjectsTianhang Cheng, Wei-Chiu Ma, Kaiyu Guan, Antonio Torralba et al.NeurIPS 2023 · 5 citations
- Can Large Language Models Understand Symbolic Graphics Programs?Zeju Qiu, Weiyang Liu, Haiwen Feng, Zhen Liu et al.ICLR 2025
Builds on5
- Free-Form Image Inpainting With Gated ConvolutionJiahui Yu, Zhe Lin, Jimei Yang, Xiaohui Shen et al.ICCV 2019 · 1,990 citations
- End-to-End Wireframe ParsingYichao Zhou, Haozhi Qi, Yi MaICCV 2019 · 190 citations
- Program-Guided Image ManipulatorsXiuming Zhang, Jiayuan Mao, Yikai Li, William T. Freeman et al.ICCV 2019 · 25 citations
- Perspective Plane Program Induction From a Single ImageYikai Li, Jiayuan Mao, Xiuming Zhang, William T. Freeman et al.CVPR 2020
- SynSin: End-to-End View Synthesis From a Single ImageOlivia Wiles, Georgia Gkioxari, Richard Szeliski, Justin JohnsonCVPR 2020
Related papers
- BoxCtrl: 3D-Aware Visual Prompting for Geometric Image EditingFeifei Wang, Shiyuan Yang, Xiaoyu Li, Jing LiaoSIGGRAPH 2026
- Panoptic Neural Fields: A Semantic Object-Aware Neural Scene RepresentationAbhijit Kundu, Kyle Genova, Xiaoqi Yin, Alireza Fathi et al.CVPR 2022 · 204 citations
- AssetField: Assets Mining and Reconfiguration in Ground Feature Plane RepresentationYuanbo Xiangli, Linning Xu, Xingang Pan, Nanxuan Zhao et al.ICCV 2023 · 13 citations
- Learning 3D Scene Priors with 2D SupervisionYinyu Nie, Angela Dai, Xiaoguang Han, Matthias NießnerCVPR 2023
- Single-View View Synthesis in the Wild with Learned Adaptive Multiplane ImagesYuxuan Han, Ruicheng Wang, Jiaolong YangSIGGRAPH 2022 · 65 citations
