Principle Component Trees and Their Persistent Homology
Ben A. Kizaric, Daniel L. Pimentel-Alarcón
Abstract
Low dimensional models like PCA are often used to simplify complex datasets by learning a single approximating subspace. This paradigm has expanded to union of subspaces models, like those learned by subspace clustering. In this paper, we present Principal Component Trees (PCTs), a graph structure that generalizes these ideas to identify mixtures of components that together describe the subspace structure of high-dimensional datasets. Each node in a PCT corresponds to a principal component of the data, and the edges between nodes indicate the components that must be mixed to produce a subspace that approximates a portion of the data. In order to construct PCTs, we propose two angle-distribution hypothesis tests to detect subspace clusters in the data. To analyze, compare, and select the best PCT model, we define two persistent homology measures that describe their shape. We show our construction yields two key properties of PCTs, namely ancestral orthogonality and non-decreasing singular values. Our main theoretical results show that learning PCTs reduces to PCA under multivariate normality, and that PCTs are efficient parameterizations of intersecting union of subspaces. Finally, we use PCTs to analyze neural network latent space, word embeddings, and reference image datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 234c4ac2-e370-42ea-91eb-fa6e92a22548Related papers
- The tree autoencoder model, with application to hierarchical data visualizationMiguel Á. Carreira-Perpiñán, Kuat GazizovNeurIPS 2024 · 3 citations
- The Persistent Laplacian for Data Science: Evaluating Higher-Order Persistent Spectral Representations of DataThomas Davies, Zhengchao Wan, Rubén J. Sánchez-GarcíaICML 2023 · 9 citations
- Efficient Orthogonal Multi-view Subspace ClusteringMan-Sheng Chen, Chang-Dong Wang, Dong Huang, Jian-Huang Lai et al.KDD 2022 · 102 citations
- Mapping the Multiverse of Latent RepresentationsJeremy Wayland, Corinna Coupette, Bastian RieckICML 2024 · 10 citations
- Exploring a Principled Framework for Deep Subspace ClusteringXianghan Meng, Zhiyuan Huang, Wei He, Xianbiao Qi et al.ICLR 2025
