Object segmentation from common fate: Motion energy processing enables human-like zero-shot generalization to random dot stimuli
Matthias Tangemann, Matthias Kümmerer, Matthias Bethge
Abstract
Humans excel at detecting and segmenting moving objects according to the Gestalt principle of"common fate". Remarkably, previous works have shown that human perception generalizes this principle in a zero-shot fashion to unseen textures or random dots. In this work, we seek to better understand the computational basis for this capability by evaluating a broad range of optical flow models and a neuroscience inspired motion energy model for zero-shot figure-ground segmentation of random dot stimuli. Specifically, we use the extensively validated motion energy model proposed by Simoncelli and Heeger in 1998 which is fitted to neural recordings in cortex area MT. We find that a cross section of 40 deep optical flow models trained on different datasets struggle to estimate motion patterns in random dot videos, resulting in poor figure-ground segmentation performance. Conversely, the neuroscience-inspired model significantly outperforms all optical flow models on this task. For a direct comparison to human perception, we conduct a psychophysical study using a shape identification task as a proxy to measure human segmentation performance. All state-of-the-art optical flow models fall short of human performance, but only the motion energy model matches human capability. This neuroscience-inspired model successfully addresses the lack of human-like zero-shot generalization to random dot stimuli in current computer vision models, and thus establishes a compelling link between the Gestalt psychology of human object perception and cortical motion processing in the brain. Code, models and datasets are available at https://github.com/mtangemann/motion_energy_segmentation
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Time Blindness: Why Video-Language Models Can't See What Humans Can?Ujjwal Upadhyay, Mukul Ranjan, Zhiqiang Shen, Mohamed ElhoseinyCVPR 2026 · 17 citations
- GenMatter: Perceiving Physical Objects with Generative Matter ModelsEric Li, Arijit Dasgupta, Yoni Friedman, Mathieu Huot et al.CVPR 2026
Builds on6
- Learning to Estimate Hidden Motions with Global Motion AggregationShihao Jiang, Dylan Campbell, Yao Lu, Hongdong Li et al.ICCV 2021 · 402 citations
- GMFlow: Learning Optical Flow via Global MatchingHaofei Xu, Jing Zhang, Jianfei Cai, Hamid Rezatofighi et al.CVPR 2022 · 353 citations
- Kubric: A scalable dataset generatorKlaus Greff, Francois Belletti, Lucas Beyer, Carl Doersch et al.CVPR 2022 · 183 citations
- Segmenting Moving Objects via an Object-Centric Layered RepresentationJunyu Xie, Weidi Xie, Andrew ZissermanNeurIPS 2022 · 74 citations
- Modeling Human Visual Motion Processing with Trainable Motion Energy Sensing and a Self-attention NetworkZitang Sun, Yen-Ju Chen, Yung-Hao Yang, Shin'ya NishidaNeurIPS 2023 · 10 citations
Related papers
- Bootstrapping Objectness from Videos by Relaxed Common Fate and Visual GroupingLong Lian, Zhirong Wu, Stella X. YuCVPR 2023
- Latent Noise Segmentation: How Neural Noise Leads to the Emergence of Segmentation and GroupingBen Lonnqvist, Zhengqing Wu, Michael H. HerzogICML 2024 · 5 citations
- Flow Snapshot Neurons in Action: Deep Neural Networks Generalize to Biological Motion PerceptionShuangpeng Han, Ziyu Wang, Mengmi ZhangNeurIPS 2024 · 8 citations
- Motion-Attentive Transition for Zero-Shot Video Object SegmentationTianfei Zhou, Shunzhou Wang, Yi Zhou, Yazhou Yao et al.AAAI 2020 · 210 citations
- The Emergence of Objectness: Learning Zero-shot Segmentation from VideosRuntao Liu, Zhirong Wu, Stella X. Yu, Stephen LinNeurIPS 2021 · 62 citations
