Learning to Compose Visual Relations
Nan Liu, Shuang Li, Yilun Du, Josh Tenenbaum, Antonio Torralba
Abstract
The visual world around us can be described as a structured set of objects and their associated relations. An image of a room may be conjured given only the description of the underlying objects and their associated relations. While there has been significant work on designing deep neural networks which may compose individual objects together, less work has been done on composing the individual relations between objects. A principal difficulty is that while the placement of objects is mutually independent, their relations are entangled and dependent on each other. To circumvent this issue, existing works primarily compose relations by utilizing a holistic encoder, in the form of text or graphs. In this work, we instead propose to represent each relation as an unnormalized density (an energy-based model), enabling us to compose separate relations in a factorized manner. We show that such a factorized decomposition allows the model to both generate and edit scenes that have multiple sets of relations more faithfully. We further show that decomposition enables our model to effectively understand the underlying relational scene structure. Project page at: https://composevisualrelations.github.io/ * indicates equal contribution 35th Conference on Neural Information Processing Systems (NeurIPS 2021).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e066b146-192e-4d62-89d3-922ffff8f3faCited by top-tier papers35
- Zero-Shot Text-Guided Object Generation with Dream FieldsAjay Jain, Ben Mildenhall, Jonathan T. Barron, Pieter Abbeel et al.CVPR 2022 · 361 citations
- Reduce, Reuse, Recycle: Compositional Generation with Energy-Based Diffusion Models and MCMCYilun Du, Conor Durkan, Robin Strudel, Joshua B. Tenenbaum et al.ICML 2023 · 219 citations
- Teaching CLIP to Count to TenRoni Paiss, Ariel Ephrat, Omer Tov, Shiran Zada et al.ICCV 2023 · 196 citations
- ReCLIP: A Strong Zero-Shot Baseline for Referring Expression ComprehensionSanjay Subramanian, William Merrill, Trevor Darrell, Matt Gardner et al.ACL 2022 · 172 citations
- Video Language PlanningYilun Du, Sherry Yang, Pete Florence, Fei Xia et al.ICLR 2024 · 161 citations
Builds on10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Training Generative Adversarial Networks with Limited DataTero Karras, Miika Aittala, Janne Hellsten, Samuli Laine et al.NeurIPS 2020 · 2,345 citations
- Your classifier is secretly an energy based model and you should treat it like oneWill Grathwohl, Kuan-Chieh Wang, Jörn-Henrik Jacobsen, David Duvenaud et al.ICLR 2020 · 643 citations
- Specifying Object Attributes and Relations in Interactive Scene GenerationOron Ashual, Lior WolfICCV 2019 · 190 citations
Related papers
- Compositional Image Decomposition with Diffusion ModelsJocelin Su, Nan Liu, Yanbo Wang, Joshua B. Tenenbaum et al.ICML 2024 · 16 citations
- GIRAFFE: Representing Scenes As Compositional Generative Neural Feature FieldsMichael Niemeyer, Andreas GeigerCVPR 2021
- Generative Scene Graph NetworksFei Deng, Zhuo Zhi, Donghun Lee, Sungjin AhnICLR 2021 · 10 citations
- Disentangled 3D Scene Generation with Layout LearningDave Epstein, Ben Poole, Ben Mildenhall, Alexei A. Efros et al.ICML 2024 · 39 citations
- Unsupervised Learning of Compositional Scene Representations from Multiple Unspecified ViewpointsJinyang Yuan, Bin Li, Xiangyang XueAAAI 2022 · 12 citations
