Learning predictable and robust neural representations by straightening image sequences
Xueyan Niu, Cristina Savin, Eero P. Simoncelli
Abstract
Prediction is a fundamental capability of all living organisms, and has been proposed as an objective for learning sensory representations. Recent work demonstrates that in primate visual systems, prediction is facilitated by neural representations that follow straighter temporal trajectories than their initial photoreceptor encoding, which allows for prediction by linear extrapolation. Inspired by these experimental findings, we develop a self-supervised learning (SSL) objective that explicitly quantifies and promotes straightening. We demonstrate the power of this objective in training deep feedforward neural networks on smoothly-rendered synthetic image sequences that mimic commonly-occurring properties of natural videos. The learned model contains neural embeddings that are predictive, but also factorize the geometric, photometric, and semantic attributes of objects. The representations also prove more robust to noise and adversarial attacks compared to previous SSL methods that optimize for invariance to random augmentations. Moreover, these beneficial properties can be transferred to other training procedures by using the straightening objective as a regularizer, suggesting a broader utility for straightening as a principle for robust unsupervised learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ae1d41f6-044e-4a7f-9097-da6ebe45e216Cited by top-tier papers3
- AI-Generated Video Detection via Perceptual StraighteningChristian Internò, Robert Geirhos, Markus Olhofer, Sunny Liu et al.NeurIPS 2025 · 44 citations
- Temporal Straightening for Latent PlanningYing Wang, Oumayma Bounou, Gaoyue Zhou, Randall Balestriero et al.ICML 2026 · 19 citations
- Chirality in Action: Time-Aware Video Representation Learning by Latent StraighteningPiyush Bagad, Andrew ZissermanNeurIPS 2025 · 14 citations
Builds on14
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Barlow Twins: Self-Supervised Learning via Redundancy ReductionJure Zbontar, Li Jing, Ishan Misra, Yann LeCun et al.ICML 2021 · 2,942 citations
- Data-Efficient Image Recognition with Contrastive Predictive CodingOlivier J. HénaffICML 2020 · 1,553 citations
Related papers
- Brain-like representational straightening of natural movies in robust feedforward neural networksTahereh Toosi, Elias B. IssaICLR 2023 · 4 citations
- Exploring perceptual straightness in learned visual representationsAnne Harrington, Vasha DuTell, Ayush Tewari, Mark Hamilton et al.ICLR 2023
- Contrastive-Equivariant Self-Supervised Learning Improves Alignment with Primate Visual Area ITThomas E. Yerxa, Jenelle Feather, Eero P. Simoncelli, SueYeon ChungNeurIPS 2024 · 12 citations
- A polar prediction model for learning to represent visual transformationsPierre-Étienne H. Fiquet, Eero P. SimoncelliNeurIPS 2023 · 9 citations
- Equivariant Self-Supervised Learning: Encouraging Equivariance in RepresentationsRumen Dangovski, Li Jing, Charlotte Loh, Seungwook Han et al.ICLR 2022 · 54 citations
