A polar prediction model for learning to represent visual transformations
Pierre-Étienne H. Fiquet, Eero P. Simoncelli
Abstract
All organisms make temporal predictions, and their evolutionary fitness level depends on the accuracy of these predictions. In the context of visual perception, the motions of both the observer and objects in the scene structure the dynamics of sensory signals, allowing for partial prediction of future signals based on past ones. Here, we propose a self-supervised representation-learning framework that extracts and exploits the regularities of natural videos to compute accurate predictions. We motivate the polar architecture by appealing to the Fourier shift theorem and its group-theoretic generalization, and we optimize its parameters on next-frame prediction. Through controlled experiments, we demonstrate that this approach can discover the representation of simple transformation groups acting in data. When trained on natural video datasets, our framework achieves better prediction performance than traditional motion compensation and rivals conventional deep networks, while maintaining interpretability and speed. Furthermore, the polar computations can be restructured into components resembling normalized simple and direction-selective complex cell models of primate V1 neurons. Thus, polar prediction offers a principled framework for understanding how the visual system represents sensory inputs in a form that simplifies temporal prediction.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7ff0469a-9fc9-477d-bd68-1d87095f77e1Cited by top-tier papers5
- Pre-trained Large Language Models Use Fourier Features to Compute AdditionTianyi Zhou, Deqing Fu, Vatsal Sharan, Robin JiaNeurIPS 2024 · 48 citations
- FoNE: Precise Single-Token Number Embeddings via Fourier FeaturesTianyi Zhou, Deqing Fu, Mahdi Soltanolkotabi, Robin Jia et al.ICLR 2026 · 24 citations
- Poisson Variational AutoencoderHadi Vafaii, Dekel Galor, Jacob L. YatesNeurIPS 2024 · 18 citations
- Learning predictable and robust neural representations by straightening image sequencesXueyan Niu, Cristina Savin, Eero P. SimoncelliNeurIPS 2024 · 13 citations
- Brain-like Variational InferenceHadi Vafaii, Dekel Galor, Jacob L. YatesNeurIPS 2025 · 7 citations
Builds on3
- Forecasting Sequential Data Using Consistent Koopman AutoencodersOmri Azencot, N. Benjamin Erichson, Vanessa Lin, Michael W. MahoneyICML 2020 · 203 citations
- Bispectral Neural NetworksSophia Sanborn, Christian Shewmake, Bruno A. Olshausen, Christopher J. HillarICLR 2023 · 78 citations
- Biological Learning of Irreducible Representations of Commuting TransformationsAlexander Genkin, David Lipshutz, Siavash Golkar, Tiberiu Tesileanu et al.NeurIPS 2022 · 5 citations
Related papers
- Learning V1 Simple Cells with Vector Representation of Local Content and Matrix Representation of Local MotionRuiqi Gao, Jianwen Xie, Siyuan Huang, Yufan Ren et al.AAAI 2022 · 2 citations
- VCT: A Video Compression TransformerFabian Mentzer, George Toderici, David Minnen, Sergi Caelles et al.NeurIPS 2022 · 155 citations
- Video Playback Rate Perception for Self-Supervised Spatio-Temporal Representation LearningYuan Yao, Chang Liu, Dezhao Luo, Yu Zhou et al.CVPR 2020
- Learning the Predictability of the FutureDidac Suris, Ruoshi Liu, Carl VondrickCVPR 2021
- Self-Supervised Representation Learning from Flow EquivarianceYuwen Xiong, Mengye Ren, Wenyuan Zeng, Raquel Urtasun WaabiICCV 2021 · 32 citations
