GenMatter: Perceiving Physical Objects with Generative Matter Models
Eric Li, Arijit Dasgupta, Yoni Friedman, Mathieu Huot, Vikash Mansinghka, Thomas O'Connell, William Freeman, Joshua B. Tenenbaum
摘要
Human visual perception offers valuable insights for understanding computational principles of motion-based scene interpretation. Humans robustly detect and segment moving entities that constitute independently moveable chunks of matter, whether observing sparse moving dots, textured surfaces, or naturalistic scenes. In contrast, existing computer vision systems lack a unified approach that works across these diverse settings. Inspired by principles of human perception, we propose a generative model that hierarchically groups low-level motion and appearance features into particles (small Gaussians representing local matter), and groups particles into clusters capturing coherently and independently moveable physical entities. We develop a hardware-accelerated inference algorithm based on parallelized block Gibbs sampling to recover stable particle motion and groupings. Our model operates on different kinds of inputs (random dots, stylized textures, or naturalistic RGB video), enabling it to work across settings where biological vision succeeds but existing computer vision approaches do not. We validate this unified framework across three domains: on 2D random dot kinematograms, our approach captures human object perception including graded uncertainty across ambiguous conditions; on a Gestalt-inspired dataset of camouflaged rotating objects, our approach recovers correct 3D structure from motion and thereby accurate 2D object segmentation; and on naturalistic RGB videos, our model tracks the moving 3D matter that makes up deforming objects, enabling robust object-level scene understanding. This work thus establishes a general framework for motion-based perception grounded in principles of human vision.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- HuMoR: 3D Human Motion Model for Robust Pose EstimationDavis Rempe, Tolga Birdal, Aaron Hertzmann, Jimei Yang 等ICCV 2021 · 被引用 398 次
- TAPIR: Tracking Any Point with per-frame Initialization and temporal RefinementCarl Doersch, Yi Yang, Mel Vecerík, Dilara Gokay 等ICCV 2023 · 被引用 297 次
- Tracking Everything Everywhere All at OnceQianqian Wang, Yen-Yu Chang, Ruojin Cai, Zhengqi Li 等ICCV 2023 · 被引用 238 次
- Diffusion with Forward Models: Solving Stochastic Inverse Problems Without Direct SupervisionAyush Tewari, Tianwei Yin, George Cazenavette, Semon Rezchikov 等NeurIPS 2023 · 被引用 131 次
- Watch It Move: Unsupervised Discovery of 3D Joints for Re-Posing of Articulated ObjectsAtsuhiro Noguchi, Umar Iqbal, Jonathan Tremblay, Tatsuya Harada 等CVPR 2022 · 被引用 32 次
相关 Paper
- PARTS: Unsupervised segmentation with slots, attention and independence maximizationDaniel Zoran, Rishabh Kabra, Alexander Lerchner, Danilo J. RezendeICCV 2021 · 被引用 53 次
- Compositional Video PredictionYufei Ye, Maneesh Singh, Abhinav Gupta, Shubham TulsianiICCV 2019 · 被引用 84 次
- TRACE: Learning 3D Gaussian Physical Dynamics from Multi-View VideosJinxi Li, Ziyang Song, Bo YangICCV 2025 · 被引用 3 次
- Articulation in Motion: Prior-free Part Mobility Analysis for Articulated Objects By Dynamic-Static DisentanglementHao Ai, Wenjie Chang, Jianbo Jiao, Ales Leonardis 等ICLR 2026 · 被引用 7 次
- Generative Perception of Shape and Material from Differential MotionXinran Nicole Han, Ko Nishino, Todd E. ZicklerNeurIPS 2025 · 被引用 2 次
