Tree-Wasserstein Distance for High Dimensional Data with a Latent Feature Hierarchy
Ya-Wei Eileen Lin, Ronald R. Coifman, Gal Mishne, Ronen Talmon
Abstract
Finding meaningful distances between high-dimensional data samples is an important scientific task. To this end, we propose a new tree-Wasserstein distance (TWD) for high-dimensional data with two key aspects. First, our TWD is specifically designed for data with a latent feature hierarchy, i.e., the features lie in a hierarchical space, in contrast to the usual focus on embedding samples in hyperbolic space. Second, while the conventional use of TWD is to speed up the computation of the Wasserstein distance, we use its inherent tree as a means to learn the latent feature hierarchy. The key idea of our method is to embed the features into a multi-scale hyperbolic space using diffusion geometry and then present a new tree decoding method by establishing analogies between the hyperbolic embedding and trees. We show that our TWD computed based on data observations provably recovers the TWD defined with the latent feature hierarchy and that its computation is efficient and scalable. We showcase the usefulness of the proposed TWD in applications to word-document and single-cell RNA-sequencing datasets, demonstrating its advantages over existing TWDs and methods based on pre-trained models. Recently, hyperbolic geometry (Ratcliffe et al., 1994) has gained prominence in hierarchical representation learning (Chamberlain et al., 2017; Nickel & Kiela, 2017) because the lengths of geodesic paths in hyperbolic spaces grow exponentially with the radius (Sarkar, 2011), a property that naturally mirrors the exponential growth of the number of nodes in hierarchical structures as the depth increases. Methods using hyperbolic geometry typically focus on finding a hyperbolic embedding of the samples, relying on a (partially) known graph, whose nodes represent the samples (Sala et al., 2018) . However, considering such a known hierarchical structure of the samples is fundamentally different than the problem we consider here, where we aim to find meaningful distances between data samples that incorporate the latent hierarchical structure of the features. In this paper, we introduce a new tree-Wasserstein distance (TWD) (Indyk & Thaper, 2003) for this purpose, where we model samples as distributions supported on a latent hierarchical structure. We propose a two-step approach. In the first step, we embed features into continuous hyperbolic spaces (Bowditch, 2007) utilizing diffusion geometry (Coifman & Lafon, 2006) to approximate the hidden
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b5bf7d88-f9bb-40f4-8933-d31b19c40907Cited by top-tier papers11
- Joint Hierarchical Representation Learning of Samples and Features via Informed Tree-Wasserstein DistanceYa-Wei Eileen Lin, Ronald R. Coifman, Gal Mishne, Ronen TalmonNeurIPS 2025 · 3 citations
- Tree-Sliced Entropy Partial TransportViet-Hoang Tran, Thanh Tran, Thanh T. Chu, Tam Le et al.NeurIPS 2025 · 3 citations
- Learning Eigenstructures of Unstructured Data ManifoldsRoy Velich, Arkadi Piven, David Bensaïd, Daniel Cremers et al.CVPR 2026 · 1 citation
- Supervised and Semi-Supervised Diffusion Maps with Label-Driven DiffusionHarel Mendelman, Ronen TalmonICLR 2025
- UltraTWD: Optimizing Ultrametric Trees for Tree-Wasserstein DistanceFangchen Yu, Yanzhen Chen, Jiaxing Wei, Jianfeng Mao et al.ICML 2025
Builds on16
- Faster Wasserstein Distance Estimation with the Sinkhorn DivergenceLénaïc Chizat, Pierre Roussillon, Flavien Léger, François-Xavier Vialard et al.NeurIPS 2020 · 164 citations
- Manifold Interpolating Optimal-Transport Flows for Trajectory InferenceGuillaume Huguet, Daniel Sumner Magruder, Alexander Tong, Oluwadamilola Fasina et al.NeurIPS 2022 · 126 citations
- From Trees to Continuous Embeddings and Back: Hyperbolic Hierarchical ClusteringInes Chami, Albert Gu, Vaggos Chatziafratis, Christopher RéNeurIPS 2020 · 125 citations
- Metric Flow Matching for Smooth Interpolations on the Data ManifoldKacper Kapusniak, Peter Potaptchik, Teodora Reu, Leo Zhang et al.NeurIPS 2024 · 89 citations
- Differentiating through the Fréchet MeanAaron Lou, Isay Katsman, Qingxuan Jiang, Serge J. Belongie et al.ICML 2020 · 83 citations
Related papers
- Hyperbolic Diffusion Embedding and Distance for Hierarchical Representation LearningYa-Wei Eileen Lin, Ronald R. Coifman, Gal Mishne, Ronen TalmonICML 2023 · 26 citations
- Fast unsupervised ground metric learning with tree-Wasserstein distanceKira Michaela Düsterwald, Samo Hromadka, Makoto YamadaICLR 2025
- A linear time approximation of Wasserstein distance with word embedding selectionSho Otao, Makoto YamadaEMNLP 2023 · 2 citations
- Wasserstein Wormhole: Scalable Optimal Transport Distance with TransformerDoron Haviv, Russell Zhang Kunes, Thomas Dougherty, Cassandra Burdziak et al.ICML 2024 · 15 citations
- Tree! I am no Tree! I am a low dimensional Hyperbolic EmbeddingRishi Sonthalia, Anna C. GilbertNeurIPS 2020 · 62 citations
