Low-Pass Filtering Improves Behavioral Alignment of Vision Models
Max Wolff, Thomas Klein, Evgenia Rusak, Felix A. Wichmann, Wieland Brendel
摘要
Despite their impressive performance on computer vision benchmarks, Deep Neural Networks (DNNs) still fall short of adequately modeling human visual behavior, as measured by error consistency and shape bias. Recent work hypothesized that behavioral alignment can be drastically improved through generative-rather than discriminative-classifiers, with far-reaching implications for models of human vision. Here, we instead show that the increased alignment of generative models can be largely explained by a seemingly innocuous resizing operation in the generative model which effectively acts as a low-pass filter. In a series of controlled experiments, we show that removing high-frequency spatial information from discriminative models like CLIP drastically increases their behavioral alignment. Simply blurring images at test-time-rather than training on blurred imagesachieves a new state-of-the-art score on the model-vs-human benchmark, halving the current alignment gap between DNNs and human observers. Furthermore, lowpass filters are likely optimal, which we demonstrate by directly optimizing filters for alignment. To contextualize the performance of optimal filters, we compute the frontier of all possible pareto-optimal solutions to the benchmark, which was formerly unknown. We explain our findings by observing that the frequency spectrum of optimal Gaussian filters roughly matches the spectrum of band-pass filters implemented by the human visual system. We show that the contrast sensitivity function, describing the inverse of the contrast threshold required for humans to detect a sinusoidal grating as a function of spatiotemporal frequency, is approximated well by Gaussian filters of the specific width that also maximizes error consistency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Scaling Vision Transformers to 22 Billion ParametersMostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski 等ICML 2023 · 被引用 848 次
- Partial success in closing the gap between human and machine visionRobert Geirhos, Kantharaju Narayanappa, Benjamin Mitzkus, Tizian Thieringer 等NeurIPS 2021 · 被引用 304 次
相关 Paper
- Quantifying Uncertainty in Error Consistency: Towards Reliable Behavioral Comparison of ClassifiersThomas Klein, Sascha Meyen, Wieland Brendel, Felix A. Wichmann 等NeurIPS 2025 · 被引用 2 次
- Model-Behavior Alignment under Flexible Evaluation: When the Best-Fitting Model Isn't the Right OneItamar Avitan, Tal GolanNeurIPS 2025 · 被引用 5 次
- Human alignment of neural network representationsLukas Muttenthaler, Jonas Dippel, Lorenz Linhardt, Robert A. Vandermeulen 等ICLR 2023 · 被引用 15 次
- Dimensionality Mismatch Between Brains and Artificial Neural NetworksSantiago Galella, Maren H. Wehrheim, Matthias KaschubeNeurIPS 2025
- Scaling Laws for Task-Optimized Models of the Primate Visual Ventral StreamAbdülkadir Gökce, Martin SchrimpfICML 2025
