ICML2026

Understanding Generalization from Embedding Dimension and Distributional Convergence

Junjie Yu, Zhuoli Ouyang, Haotian Deng, Chen Wei, Wenxiao Ma, Jianyu Zhang, Zihan Deng, Quanying Liu

Abstract

Deep neural networks often generalize well despite heavy over-parameterization, challenging classical parameter-based analyses. We study generalization from a representation-centric perspective and analyze how the geometry of learned embeddings is associated with generalization performance for a fixed trained model. We derive a post-hoc generalization bound that relates the gap between population risk and held-out empirical risk to two factors: (i) the intrinsic dimension of the embedding, which determines the convergence rate of the empirical embedding distribution to its population counterpart in Wasserstein distance, and (ii) the sensitivity of the downstream mapping from embeddings to predictions, characterized by Lipschitz constants. Together, these provide a post-hoc explanation of generalization for trained models. At the final embedding layer, architectural sensitivity disappears and the bound is dominated by embedding dimension, explaining its strong empirical correlation with generalization performance. Experiments across architectures and datasets validate the theory and demonstrate the utility of embedding-based diagnostics.