Residual Connections Harm Generative Representation Learning
Xiao Zhang, Ruoxi Jiang, William Gao, Rebecca Willett, Michael Maire
摘要
We show that introducing a weighting factor to reduce the influence of identity shortcuts in residual networks significantly enhances semantic feature learning in generative representation learning frameworks, such as masked autoencoders (MAEs) and diffusion models. Our modification notably improves feature quality, raising ImageNet-1K K-Nearest Neighbor accuracy from 27.4% to 63.9% and linear probing accuracy from 67.8% to 72.7% for MAEs with a ViT-B/16 backbone, while also enhancing generation quality in diffusion models. This significant gap suggests that, while residual connection structure serves an essential role in facilitating gradient propagation, it may have a harmful side effect of reducing capacity for abstract learning by virtue of injecting an echo of shallower representations into deeper layers. We ameliorate this downside via a fixed formula for monotonically decreasing the contribution of identity connections as layer depth increases. Our design promotes the gradual development of feature abstractions, without impacting network trainability. Analyzing the representations learned by our modified residual networks, we find correlation between low effective feature rank and downstream task performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Auto-Compressing NetworksEvangelos Dorovatas, Georgios Paraskevopoulos, Alexandros PotamianosNeurIPS 2025 · 被引用 4 次
- Relighting as a Probe of Visual Priors via Augmented Latent IntrinsicsXiaoyan Xing, Xiao Zhang, Sezer Karaoglu, Theo Gevers 等ICML 2026
- Uni-Encoder Meets Multi-Encoders: Representation Before Fusion for Brain Tumor Segmentation with Missing ModalitiesPeibo Song, Xiaotian Xue, Jinshuo Zhang, Zihao Wang 等CVPR 2026
它引用的顶会 Paper29
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
相关 Paper
- Masked Autoencoders Are Effective Tokenizers for Diffusion ModelsHao Chen, Yujin Han, Fangyi Chen, Xiang Li 等ICML 2025
- Masked Image Residual Learning for Scaling Deeper Vision TransformersGuoxi Huang, Hongtao Fu, Adrian G. BorsNeurIPS 2023 · 被引用 10 次
- Not All Parameters Matter: Masking Diffusion Models for Enhancing Generation AbilityLei Wang, Senmao Li, Fei Yang, Jianye Wang 等CVPR 2025
- Deep Compression Autoencoder for Efficient High-Resolution Diffusion ModelsJunyu Chen, Han Cai, Junsong Chen, Enze Xie 等ICLR 2025 · 被引用 2 次
- MAGE: MAsked Generative Encoder to Unify Representation Learning and Image SynthesisTianhong Li, Huiwen Chang, Shlok Kumar Mishra, Han Zhang 等CVPR 2023
