Exploring perceptual straightness in learned visual representations
Anne Harrington, Vasha DuTell, Ayush Tewari, Mark Hamilton, Simon Stent, Ruth Rosenholtz, William T. Freeman
摘要
Humans have been shown to use a ''straightened'' encoding to represent the natural visual world as it evolves in time (Henaff et al. 2019). In the context of discrete video sequences, ''straightened'' means that changes between frames follow a more linear path in representation space at progressively deeper levels of processing. While deep convolutional networks are often proposed as models of human visual processing, many do not straighten natural videos. In this paper, we explore the relationship between network architecture, differing types of robustness, biologically-inspired filtering mechanisms, and representational straightness in response to time-varying input; we identify strengths and limitations of straightness as a useful way of evaluating neural network representations. We find that (1) adversarial training leads to straighter representations in both CNN and transformer-based architectures but (2) this effect is task-dependent, not generalizing to tasks such as segmentation and frame-prediction, where straight representations are not favorable for predictions; and nor to other types of robustness. In addition, (3) straighter representations impart temporal stability to class predictions, even for out-of-distribution data. Finally, (4) biologically-inspired elements increase straightness in the early stages of a network, but do not guarantee increased straightness in downstream layers of CNNs. We show that straightness is an easily computed measure of representational robustness and stability, as well as a hallmark of human representations with benefits for computer vision models.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper5
- AI-Generated Video Detection via Perceptual StraighteningChristian Internò, Robert Geirhos, Markus Olhofer, Sunny Liu 等NeurIPS 2025 · 被引用 44 次
- Temporal Straightening for Latent PlanningYing Wang, Oumayma Bounou, Gaoyue Zhou, Randall Balestriero 等ICML 2026 · 被引用 19 次
- Chirality in Action: Time-Aware Video Representation Learning by Latent StraighteningPiyush Bagad, Andrew ZissermanNeurIPS 2025 · 被引用 14 次
- Learning predictable and robust neural representations by straightening image sequencesXueyan Niu, Cristina Savin, Eero P. SimoncelliNeurIPS 2024 · 被引用 13 次
- COCO-Periph: Bridging the Gap Between Human and Machine Perception in the PeripheryAnne Harrington, Vasha DuTell, Mark Hamilton, Ayush Tewari 等ICLR 2024 · 被引用 6 次
相关 Paper
- Brain-like representational straightening of natural movies in robust feedforward neural networksTahereh Toosi, Elias B. IssaICLR 2023 · 被引用 4 次
- Instant Video Models: Universal Adapters for Stabilizing Image-Based NetworksMatthew Dutson, Nathan Labiosa, Yin Li, Mohit GuptaNeurIPS 2025
- Stretching Beyond the Obvious: A Gradient-Free Framework to Unveil the Hidden Landscape of Visual InvarianceLorenzo Tausani, Paolo Muratore, Morgan Bruce Talbot, Giacomo Amerio 等ICLR 2026
- VCT: A Video Compression TransformerFabian Mentzer, George Toderici, David Minnen, Sergi Caelles 等NeurIPS 2022 · 被引用 155 次
- Large language models implicitly learn to straighten neural sentence trajectories to construct a predictive representation of natural languageEghbal A. Hosseini, Evelina FedorenkoNeurIPS 2023 · 被引用 41 次
