Learning Physical Graph Representations from Visual Scenes
Daniel Bear, Chaofei Fan, Damian Mrowca, Yunzhu Li, Seth Alter, Aran Nayebi, Jeremy Schwartz, Li Fei-Fei, Jiajun Wu, Josh Tenenbaum, Daniel L. K. Yamins
摘要
Convolutional Neural Networks (CNNs) have proved exceptional at learning representations for visual object categorization. However, CNNs do not explicitly encode objects, parts, and their physical properties, which has limited CNNs' success on tasks that require structured understanding of visual scenes. To overcome these limitations, we introduce the idea of "Physical Scene Graphs" (PSGs), which represent scenes as hierarchical graphs, with nodes in the hierarchy corresponding intuitively to object parts at different scales, and edges to physical connections between parts. Bound to each node is a vector of latent attributes that intuitively represent object properties such as surface shape and texture. We also describe PSGNet, a network architecture that learns to extract PSGs by reconstructing scenes through a PSG-structured bottleneck. PSGNet augments standard CNNs by including: recurrent feedback connections to combine low and high-level image information; graph pooling and vectorization operations that convert spatially-uniform feature maps into object-centric graph structures; and perceptual grouping principles to encourage the identification of meaningful scene elements. We show that PSGNet outperforms alternative self-supervised scene representation algorithms at scene segmentation tasks, especially on complex real-world images, and generalizes well to unseen object types and scene arrangements. PSGNet is also able learn from physical motion, enhancing scene estimates even for static images. We present a series of ablation studies illustrating the importance of each component of the PS-GNet architecture, analyses showing that learned latent attributes capture intuitive scene properties, and illustrate the use of PSGs for compositional scene inference.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- Conditional Object-Centric Learning from VideoThomas Kipf, Gamaleldin Fathy Elsayed, Aravindh Mahendran, Austin Stone 等ICLR 2022 · 被引用 290 次
- SAVi++: Towards End-to-End Object-Centric Learning from Real-World VideosGamaleldin F. Elsayed, Aravindh Mahendran, Sjoerd van Steenkiste, Klaus Greff 等NeurIPS 2022 · 被引用 218 次
- GENESIS-V2: Inferring Unordered Object Representations without Iterative RefinementMartin Engelcke, Oiwi Parker Jones, Ingmar PosnerNeurIPS 2021 · 被引用 143 次
- Dynamic Visual Reasoning by Learning Differentiable Physics Models from Video and LanguageMingyu Ding, Zhenfang Chen, Tao Du, Ping Luo 等NeurIPS 2021 · 被引用 90 次
- Unsupervised Part Discovery from Contrastive ReconstructionSubhabrata Choudhury, Iro Laina, Christian Rupprecht, Andrea VedaldiNeurIPS 2021 · 被引用 74 次
它引用的顶会 Paper6
- Learning to Simulate Complex Physics with Graph NetworksAlvaro Sanchez-Gonzalez, Jonathan Godwin, Tobias Pfaff, Rex Ying 等ICML 2020 · 被引用 1,439 次
- GENESIS: Generative Scene Inference and Sampling with Object-Centric Latent RepresentationsMartin Engelcke, Adam R. Kosiorek, Oiwi Parker Jones, Ingmar PosnerICLR 2020 · 被引用 334 次
- Visual Grounding of Learned Physical ModelsYunzhu Li, Toru Lin, Kexin Yi, Daniel Bear 等ICML 2020 · 被引用 88 次
- Disentangling neural mechanisms for perceptual groupingJunkyung Kim, Drew Linsley, Kalpit Thakkar, Thomas SerreICLR 2020 · 被引用 61 次
- Stable and expressive recurrent vision modelsDrew Linsley, Alekh Karkada Ashok, Lakshmi Narasimhan Govindarajan, Rex G. Liu 等NeurIPS 2020 · 被引用 56 次
相关 Paper
- Generative Scene Graph NetworksFei Deng, Zhuo Zhi, Donghun Lee, Sungjin AhnICLR 2021 · 被引用 10 次
- Unconditional Scene Graph GenerationSarthak Garg, Helisa Dhamo, Azade Farshad, Sabrina Musatian 等ICCV 2021 · 被引用 30 次
- Learning 3D Semantic Scene Graphs From 3D Indoor ReconstructionsJohanna Wald, Helisa Dhamo, Nassir Navab, Federico TombariCVPR 2020
- Self-Supervised Learning of Hybrid Part-Aware 3D Representations of 2D Gaussians and SuperquadricsZhirui Gao, Renjiao Yi, Yuhang Huang, Wei Chen 等ICCV 2025 · 被引用 2 次
- Unsupervised Discovery of 3D Physical Objects from VideoYilun Du, Kevin A. Smith, Tomer D. Ullman, Joshua B. Tenenbaum 等ICLR 2021 · 被引用 11 次
