ICML2026
How High is ‘High’? Rethinking the Roles of Dimensionality in Topological Data Analysis and Manifold Learning
Hannah Sansford, Nick Whiteley, Patrick Rubin-Delanchy
被引用 1 次
摘要
High-dimensionality of data is often regarded as a fundamental statistical impediment in Machine Learning and AI. The purpose of this paper is to clarify, on the contrary, when and how high-dimensionality may be beneficial. In the setting of a general random function model of data we delineate between three notions of dimensionality: effective dimension , measuring total variability across feature directions; correlation rank , measuring functional complexity across samples; and latent intrinsic dimension of manifold structure hidden in data. Via a generalized Hanson-Wright inequality, we show that increasing drives a blessing of dimensionality phenomenon, whereby data dot-products concentrate about their expectations. In turn, we show that, under mild continuity assumptions (ensuring that features bring additional information as dimension grows), persistence diagrams recover latent homology when as . Informed by our theory, we revisit the ground-breaking neuroscience discovery of toroidal structure in grid-cell activity made by Gardner et al. (2022): our findings provide the first empirical evidence that this structure is isometric to a flat torus model of physical space, suggesting that grid cell activity conveys a geometrically faithful representation of the real world.