Lune

ICML2026Top-tier venue

Understanding Generalization from Embedding Dimension and Distributional Convergence

Junjie Yu, Zhuoli Ouyang, Haotian Deng, Chen Wei, Wenxiao Ma, Jianyu Zhang, Zihan Deng, Quanying Liu

2026Year

Abstract

Deep neural networks often generalize well despite heavy over-parameterization, challenging classical parameter-based analyses. We study generalization from a representation-centric perspective and analyze how the geometry of learned embeddings is associated with generalization performance for a fixed trained model. We derive a post-hoc generalization bound that relates the gap between population risk and held-out empirical risk to two factors: (i) the intrinsic dimension of the embedding, which determines the convergence rate of the empirical embedding distribution to its population counterpart in Wasserstein distance, and (ii) the sensitivity of the downstream mapping from embeddings to predictions, characterized by Lipschitz constants. Together, these provide a post-hoc explanation of generalization for trained models. At the final embedding layer, architectural sensitivity disappears and the bound is dominated by embedding dimension, explaining its strong empirical correlation with generalization performance. Experiments across architectures and datasets validate the theory and demonstrate the utility of embedding-based diagnostics.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 302030ee-b46d-4df9-9df5-65eecca665eb

Builds on2

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines