Transformable Bottleneck Networks
Kyle Olszewski, Sergey Tulyakov, Oliver J. Woodford, Hao Li, Linjie Luo
Abstract
We propose a novel approach to performing fine-grained 3D manipulation of image content via a convolutional neural network, which we call the Transformable Bottleneck Network (TBN). It applies given spatial transformations directly to a volumetric bottleneck within our encoderbottleneck-decoder architecture. Multi-view supervision encourages the network to learn to spatially disentangle the feature space within the bottleneck. The resulting spatial structure can be manipulated with arbitrary spatial transformations. We demonstrate the efficacy of TBNs for novel view synthesis, achieving state-of-the-art results on a challenging benchmark. We demonstrate that the bottlenecks produced by networks trained for this task contain meaningful spatial structure that allows us to intuitively perform a variety of image manipulations in 3D, well beyond the rigid transformations seen during training. These manipulations include non-uniform scaling, non-rigid warping, and combining content from different images. Finally, we extract explicit 3D structure from the bottleneck, performing impressive 3D reconstruction from a single input image. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 370b7319-d80a-47c0-b25e-9ba6e4c3f7edCited by top-tier papers17
- Putting NeRF on a Diet: Semantically Consistent Few-Shot View SynthesisAjay Jain, Matthew Tancik, Pieter AbbeelICCV 2021 · 615 citations
- Make-It-3D: High-Fidelity 3D Creation from A Single Image with Diffusion PriorJunshu Tang, Tengfei Wang, Bo Zhang, Ting Zhang et al.ICCV 2023 · 405 citations
- Neural Human Performer: Learning Generalizable Radiance Fields for Human Performance RenderingYoungjoong Kwon, Dahun Kim, Duygu Ceylan, Henry FuchsNeurIPS 2021 · 224 citations
- Unconstrained Scene Generation with Locally Conditioned Radiance FieldsTerrance DeVries, Miguel Ángel Bautista, Nitish Srivastava, Graham W. Taylor et al.ICCV 2021 · 169 citations
- Equivariant Neural RenderingEmilien Dupont, Miguel Bautista Martin, Alex Colburn, Aditya Sankar et al.ICML 2020 · 69 citations
Related papers
- Flow Guided Transformable Bottleneck Networks for Motion RetargetingJian Ren, Menglei Chai, Oliver J. Woodford, Kyle Olszewski et al.CVPR 2021
- Towards Unsupervised Learning of Generative Models for 3D Controllable Image SynthesisYiyi Liao, Katja Schwarz, Lars M. Mescheder, Andreas GeigerCVPR 2020
- Disentangled3D: Learning a 3D Generative Model with Disentangled Geometry and Appearance from Monocular ImagesAyush Tewari, Mallikarjun B. R., Xingang Pan, Ohad Fried et al.CVPR 2022 · 35 citations
- Geometry-Free View Synthesis: Transformers and no 3D PriorsRobin Rombach, Patrick Esser, Björn OmmerICCV 2021 · 115 citations
- CustomNet: Object Customization with Variable-Viewpoints in Text-to-Image Diffusion ModelsZiyang Yuan, Mingdeng Cao, Xintao Wang, Zhongang Qi et al.ACM MM 2024 · 10 citations
