Modeling Human Visual Motion Processing with Trainable Motion Energy Sensing and a Self-attention Network
Zitang Sun, Yen-Ju Chen, Yung-Hao Yang, Shin'ya Nishida
Abstract
Visual motion processing is essential for humans to perceive and interact with dynamic environments. Despite extensive research in cognitive neuroscience, imagecomputable models that can extract informative motion flow from natural scenes in a manner consistent with human visual processing have yet to be established. Meanwhile, recent advancements in computer vision (CV), propelled by deep learning, have led to significant progress in optical flow estimation, a task closely related to motion perception. Here we propose an image-computable model of human motion perception by bridging the gap between biological and CV models. Specifically, we introduce a novel two-stages approach that combines trainable motion energy sensing with a recurrent self-attention network for adaptive motion integration and segregation. This model architecture aims to capture the computations in V1-MT, the core structure for motion perception in the biological visual system, while providing the ability to derive informative motion flow for a wide range of stimuli, including complex natural scenes. In silico neurophysiology reveals that our model's unit responses are similar to mammalian neural recordings regarding motion pooling and speed tuning. The proposed model can also replicate human responses to a range of stimuli examined in past psychophysical studies. The experimental results on the Sintel benchmark demonstrate that our model predicts human responses better than the ground truth, whereas the state-of-the-art CV models show the opposite. Our study provides a computational architecture consistent with human visual motion processing, although the physiological correspondence may not be exact. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 580435e7-287b-4614-a0de-20ff1e63a68fCited by top-tier papers3
- Object segmentation from common fate: Motion energy processing enables human-like zero-shot generalization to random dot stimuliMatthias Tangemann, Matthias Kümmerer, Matthias BethgeNeurIPS 2024 · 6 citations
- Deterministic-to-Stochastic Diverse Latent Feature Mapping for Human Motion SynthesisHua Yu, Weiming Liu, Gui Xu, Yaqing Hou et al.CVPR 2025
- HuPerFlow: A Comprehensive Benchmark for Human vs. Machine Motion Estimation ComparisonYung-Hao Yang, Zitang Sun, Taiki Fukiage, Shin'ya NishidaCVPR 2025
Builds on6
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- GMFlow: Learning Optical Flow via Global MatchingHaofei Xu, Jing Zhang, Jianfei Cai, Hamid Rezatofighi et al.CVPR 2022 · 353 citations
- Learning Optical Flow with Adaptive Graph ReasoningAo Luo, Fan Yang, Kunming Luo, Xin Li et al.AAAI 2022 · 73 citations
- Your head is there to move you around: Goal-driven models of the primate dorsal pathwayPatrick J. Mineault, Shahab Bakhtiari, Blake A. Richards, Christopher C. PackNeurIPS 2021 · 61 citations
- SMURF: Self-Teaching Multi-Frame Unsupervised RAFT With Full-Image WarpingAustin Stone, Daniel Maurer, Alper Ayvaci, Anelia Angelova et al.CVPR 2021
Related papers
- Flow Snapshot Neurons in Action: Deep Neural Networks Generalize to Biological Motion PerceptionShuangpeng Han, Ziyu Wang, Mengmi ZhangNeurIPS 2024 · 8 citations
- SKFlow: Learning Optical Flow with Super KernelsShangkun Sun, Yuanqi Chen, Yu Zhu, Guodong Guo et al.NeurIPS 2022 · 97 citations
- RIRNet: Recurrent-In-Recurrent Network for Video Quality AssessmentPengfei Chen, Leida Li, Lei Ma, Jinjian Wu et al.ACM MM 2020 · 94 citations
- SimulFlow: Simultaneously Extracting Feature and Identifying Target for Unsupervised Video Object SegmentationLingyi Hong, Wei Zhang, Shuyong Gao, Hong Lu et al.ACM MM 2023 · 14 citations
- Generalizable Implicit Motion Modeling for Video Frame InterpolationZujin Guo, Wei Li, Chen Change LoyNeurIPS 2024 · 24 citations
