Local Intrinsic Dimensional Entropy
Rohan Ghosh, Mehul Motani
Abstract
Most entropy measures depend on the spread of the probability distribution over the sample space |X|, and the maximum entropy achievable scales proportionately with the sample space cardinality |X|. For a finite |X|, this yields robust entropy measures which satisfy many important properties, such as invariance to bijections, while the same is not true for continuous spaces (where |X|=infinity). Furthermore, since R and R^d (d in Z+) have the same cardinality (from Cantor's correspondence argument), cardinality-dependent entropy measures cannot encode the data dimensionality. In this work, we question the role of cardinality and distribution spread in defining entropy measures for continuous spaces, which can undergo multiple rounds of transformations and distortions, e.g., in neural networks. We find that the average value of the local intrinsic dimension of a distribution, denoted as ID-Entropy, can serve as a robust entropy measure for continuous spaces, while capturing the data dimensionality. We find that ID-Entropy satisfies many desirable properties and can be extended to conditional entropy, joint entropy and mutual-information variants. ID-Entropy also yields new information bottleneck principles and also links to causality. In the context of deep learning, for feedforward architectures, we show, theoretically and empirically, that the ID-Entropy of a hidden layer directly controls the generalization gap for both classifiers and auto-encoders, when the target function is Lipschitz continuous. Our work primarily shows that, for continuous spaces, taking a structural rather than a statistical approach yields entropy measures which preserve intrinsic data dimensionality, while being relevant for studying various architectures.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 98dbfd14-e5e3-4dc0-9feb-aa5b7c30caa1Cited by top-tier papers4
- Deep Regression Representation Learning with TopologyShihao Zhang, Kenji Kawaguchi, Angela YaoICML 2024 · 4 citations
- Intrinsic Dimension Correlation: uncovering nonlinear connections in multimodal representationsLorenzo Basile, Santiago Acevedo, Luca Bortolussi, Fabio Anselmi et al.ICLR 2025 · 1 citation
- CuMPerLay: Learning Cubical Multiparameter Persistence VectorizationsCaner Korkmaz, Brighton Nuwagira, Baris Coskunuzer, Tolga BirdalICCV 2025 · 1 citation
- Tab-PET: Graph-Based Positional Encodings for Tabular TransformersYunze Leng, Rohan Ghosh, Mehul MotaniAAAI 2026
Builds on3
- On the geometry of generalization and memorization in deep neural networksCory Stephenson, Suchismita Padhy, Abhinav Ganesh, Yue Hui et al.ICLR 2021 · 95 citations
- Intrinsic Dimension, Persistent Homology and Generalization in Neural NetworksTolga Birdal, Aaron Lou, Leonidas J. Guibas, Umut SimsekliNeurIPS 2021 · 94 citations
- The Effect of the Intrinsic Dimension on the Generalization of Quadratic ClassifiersFabian Latorre, Leello Tadesse Dadi, Paul Rolland, Volkan CevherNeurIPS 2021 · 13 citations
Related papers
- How Does Information Bottleneck Help Deep Learning?Kenji Kawaguchi, Zhun Deng, Xu Ji, Jiaoyang HuangICML 2023 · 117 citations
- Cauchy-Schwarz Divergence Information Bottleneck for RegressionShujian Yu, Xi Yu, Sigurd Løkse, Robert Jenssen et al.ICLR 2024 · 16 citations
- Information-Theoretic Generalization Bounds for VAEs: A Role of Encoder and Latent VariableFutoshi Futami, Masahiro FujisawaICML 2026
- Information Bottleneck: Exact Analysis of (Quantized) Neural NetworksStephan Sloth Lorenzen, Christian Igel, Mads NielsenICLR 2022 · 24 citations
- Learning Diverse and Discriminative Representations via the Principle of Maximal Coding Rate ReductionYaodong Yu, Kwan Ho Ryan Chan, Chong You, Chaobing Song et al.NeurIPS 2020 · 265 citations
