On the Effect of Misspecifying the Embedding Dimension in Low-rank Network Models
Roddy Taing, Keith Levin
摘要
As network data has become ubiquitous in the sciences, there has been growing interest in network models whose structure is driven by latent node-level variables in a (typically low-dimensional) latent geometric space. These "latent positions" are often estimated via embeddings, whereby the nodes of a network are mapped to points in Euclidean space so that "similar" nodes are mapped to nearby points. Under certain model assumptions, these embeddings are consistent estimates of the latent positions, but most such results require that the embedding dimension be chosen correctly, typically equal to the dimension of the latent space. Methods for estimating this correct embedding dimension have been studied extensive in recent years, but there has been little work to date characterizing the behavior of embeddings when this embedding dimension is misspecified. In this work, we provide theoretical descriptions of the effects of misspecifying the embedding dimension of the adjacency spectral embedding under the random dot product graph, a class of latent space network models that includes a number of widely-used network models as special cases, including the stochastic blockmodel. We consider both the case in which the dimension is chosen too small, where we prove estimation error lower-bounds, and the case where the dimension is chosen too large, where we show that consistency still holds, albeit at a slower rate than when the embedding dimension is chosen correctly. A range of synthetic data experiments support our theoretical results. Our main technical result, which may be of independent interest, is a generalization of earlier work in random matrix theory, showing that all non-signal eigenvectors of a low-rank matrix subject to additive noise are delocalized.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- Manifold structure in graph embeddingsPatrick Rubin-DelanchyNeurIPS 2020 · 被引用 29 次
- Coherence-free Entrywise Estimation of Eigenvectors in Low-rank Signal-plus-noise Matrix ModelsHao Yan, Keith LevinNeurIPS 2024 · 被引用 2 次
- How High is ‘High’? Rethinking the Roles of Dimensionality in Topological Data Analysis and Manifold LearningHannah Sansford, Nick Whiteley, Patrick Rubin-DelanchyICML 2026 · 被引用 1 次
相关 Paper
- Robustness of Community Detection to Random Geometric PerturbationsSandrine Péché, Vianney PerchetNeurIPS 2020 · 被引用 7 次
- Maximum Likelihood Embedding of Logistic Random Dot Product GraphsLuke J. O'Connor, Muriel Médard, Soheil FeiziAAAI 2020 · 被引用 8 次
- On the Optimization Trajectory of DeepWalk EmbeddingsChristopher Harker, Aditya BhaskaraICML 2026
- Community Detection Guarantees using Embeddings Learned by Node2VecAndrew Davison, S. Carlyle Morgan, Owen G. WardNeurIPS 2024 · 被引用 3 次
- Node Similarities under Random Projections: Limits and Pathological CasesTvrtko Tadic, Cassiano O. Becker, Jennifer NevilleICLR 2025
