Modeling Human Visual Motion Processing with Trainable Motion Energy Sensing and a Self-attention Network
Zitang Sun, Yen-Ju Chen, Yung-Hao Yang, Shin'ya Nishida
摘要
Visual motion processing is essential for humans to perceive and interact with dynamic environments. Despite extensive research in cognitive neuroscience, imagecomputable models that can extract informative motion flow from natural scenes in a manner consistent with human visual processing have yet to be established. Meanwhile, recent advancements in computer vision (CV), propelled by deep learning, have led to significant progress in optical flow estimation, a task closely related to motion perception. Here we propose an image-computable model of human motion perception by bridging the gap between biological and CV models. Specifically, we introduce a novel two-stages approach that combines trainable motion energy sensing with a recurrent self-attention network for adaptive motion integration and segregation. This model architecture aims to capture the computations in V1-MT, the core structure for motion perception in the biological visual system, while providing the ability to derive informative motion flow for a wide range of stimuli, including complex natural scenes. In silico neurophysiology reveals that our model's unit responses are similar to mammalian neural recordings regarding motion pooling and speed tuning. The proposed model can also replicate human responses to a range of stimuli examined in past psychophysical studies. The experimental results on the Sintel benchmark demonstrate that our model predicts human responses better than the ground truth, whereas the state-of-the-art CV models show the opposite. Our study provides a computational architecture consistent with human visual motion processing, although the physiological correspondence may not be exact. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Object segmentation from common fate: Motion energy processing enables human-like zero-shot generalization to random dot stimuliMatthias Tangemann, Matthias Kümmerer, Matthias BethgeNeurIPS 2024 · 被引用 6 次
- Deterministic-to-Stochastic Diverse Latent Feature Mapping for Human Motion SynthesisHua Yu, Weiming Liu, Gui Xu, Yaqing Hou 等CVPR 2025
- HuPerFlow: A Comprehensive Benchmark for Human vs. Machine Motion Estimation ComparisonYung-Hao Yang, Zitang Sun, Taiki Fukiage, Shin'ya NishidaCVPR 2025
它引用的顶会 Paper6
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- GMFlow: Learning Optical Flow via Global MatchingHaofei Xu, Jing Zhang, Jianfei Cai, Hamid Rezatofighi 等CVPR 2022 · 被引用 353 次
- Learning Optical Flow with Adaptive Graph ReasoningAo Luo, Fan Yang, Kunming Luo, Xin Li 等AAAI 2022 · 被引用 73 次
- Your head is there to move you around: Goal-driven models of the primate dorsal pathwayPatrick J. Mineault, Shahab Bakhtiari, Blake A. Richards, Christopher C. PackNeurIPS 2021 · 被引用 61 次
- SMURF: Self-Teaching Multi-Frame Unsupervised RAFT With Full-Image WarpingAustin Stone, Daniel Maurer, Alper Ayvaci, Anelia Angelova 等CVPR 2021
相关 Paper
- Flow Snapshot Neurons in Action: Deep Neural Networks Generalize to Biological Motion PerceptionShuangpeng Han, Ziyu Wang, Mengmi ZhangNeurIPS 2024 · 被引用 8 次
- SKFlow: Learning Optical Flow with Super KernelsShangkun Sun, Yuanqi Chen, Yu Zhu, Guodong Guo 等NeurIPS 2022 · 被引用 97 次
- RIRNet: Recurrent-In-Recurrent Network for Video Quality AssessmentPengfei Chen, Leida Li, Lei Ma, Jinjian Wu 等ACM MM 2020 · 被引用 94 次
- SimulFlow: Simultaneously Extracting Feature and Identifying Target for Unsupervised Video Object SegmentationLingyi Hong, Wei Zhang, Shuyong Gao, Hong Lu 等ACM MM 2023 · 被引用 14 次
- Generalizable Implicit Motion Modeling for Video Frame InterpolationZujin Guo, Wei Li, Chen Change LoyNeurIPS 2024 · 被引用 24 次
