From -SNE to UMAP with contrastive learning
Sebastian Damrich, Jan Niklas Böhm, Fred A. Hamprecht, Dmitry Kobak
摘要
Neighbor embedding methods -SNE and UMAP are the de facto standard for visualizing high-dimensional datasets. Motivated from entirely different viewpoints, their loss functions appear to be unrelated. In practice, they yield strongly differing embeddings and can suggest conflicting interpretations of the same data. The fundamental reasons for this and, more generally, the exact relationship between -SNE and UMAP have remained unclear. In this work, we uncover their conceptual connection via a new insight into contrastive learning methods. Noise-contrastive estimation can be used to optimize -SNE, while UMAP relies on negative sampling, another contrastive method. We find the precise relationship between these two contrastive methods and provide a mathematical characterization of the distortion introduced by negative sampling. Visually, this distortion results in UMAP generating more compact embeddings with tighter clusters compared to -SNE. We exploit this new conceptual connection to propose and implement a generalization of negative sampling, allowing us to interpolate between (and even extrapolate beyond) -SNE and UMAP and their respective embeddings. Moving along this spectrum of embeddings leads to a trade-off between discrete / local and continuous / global structures, mitigating the risk of over-interpreting ostensible features of any single embedding. We provide a PyTorch implementation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Improving neural network representations using human similarity judgmentsLukas Muttenthaler, Lorenz Linhardt, Jonas Dippel, Robert A. Vandermeulen 等NeurIPS 2023 · 被引用 61 次
- Dimension Reduction with Locally Adjusted GraphsYingfan Wang, Yiyang Sun, Haiyang Huang, Cynthia RudinAAAI 2025 · 被引用 10 次
- Navigating the Effect of Parametrization for Dimensionality ReductionHaiyang Huang, Yingfan Wang, Cynthia RudinNeurIPS 2024 · 被引用 8 次
- Random Forest Autoencoders for Guided Representation LearningAdrien Aumon, Shuang Ni, Myriam Lizotte, Guy Wolf 等NeurIPS 2025 · 被引用 7 次
- Unsupervised visualization of image datasets using contrastive learningJan Niklas Böhm, Philipp Berens, Dmitry KobakICLR 2023 · 被引用 6 次
它引用的顶会 Paper10
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal 等NeurIPS 2020 · 被引用 5,249 次
- Contrastive and Non-Contrastive Self-Supervised Learning Recover Global and Local Spectral Embedding MethodsRandall Balestriero, Yann LeCunNeurIPS 2022 · 被引用 189 次
- On UMAP's True Loss FunctionSebastian Damrich, Fred A. HamprechtNeurIPS 2021 · 被引用 58 次
- Understanding Negative Samples in Instance Discriminative Self-supervised Representation LearningKento Nozawa, Issei SatoNeurIPS 2021 · 被引用 56 次
相关 Paper
- Interactive Visual Cluster Analysis by Contrastive Dimensionality ReductionJiazhi Xia, Linquan Huang, Weixing Lin, Xin Zhao 等IEEE VIS 2022 · 被引用 38 次
- Your Contrastive Learning Is Secretly Doing Stochastic Neighbor EmbeddingTianyang Hu, Zhili Liu, Fengwei Zhou, Wenjia Wang 等ICLR 2023 · 被引用 3 次
- SpaceMAP: Visualizing High-Dimensional Data by Space ExpansionXinrui Zu, Qian TaoICML 2022 · 被引用 12 次
- Joint t-SNE for Comparable Projections of Multiple High-Dimensional DatasetsYinqiao Wang, Lu Chen, Jaemin Jo, Yunhai WangIEEE VIS 2021 · 被引用 33 次
- Federated t-SNE and UMAP for Distributed Data VisualizationDong Qiao, Xinxian Ma, Jicong FanAAAI 2025 · 被引用 3 次
