AutoDV: An End-to-End Deep Learning Model for High-Dimensional Data Visualization
Wei Dai, Jicong Fan
Abstract
High-dimensional data visualization (HDV) plays an important role in data science and engineering applications. Traditional HDV methods, such as Autoencoder and t-SNE, require hyper-parameter tuning and iterative optimization on every dataset and cannot effectively utilize the knowledge from historical low-dimension representation, which lowers the efficiency, convenience, and accuracy in real applications. In this paper, we present AutoDV, an end-to-end deep learning model, for high-dimensional data visualization. AutoDV is built upon a graph transformer network and an invariant loss function and is trained on a number of diverse datasets converted into multi-weight graphs. Given a new dataset, AutoDV outputs the 2D or 3D embeddings of all data points directly. AutoDV has the following merits: 1) There is no hyper-parameter selection during the data visualization stage; 2) The end-to-end model avoids re-training or iterative optimization when visualizing data; 3) The input dataset can have any number of features and can be from any domain. Our experiments show that AutoDV can successfully generalize to unseen datasets without retraining with 89.37% precision of t-SNE and 91.05% precision of UMAP on the unseen CIFAR10 datasets. Compared with existing parametric data visualization deep models, our method obtains a significant improvement with 86.65% precision gain. AutoDV can perform even better than t-SNE and UMAP on gene and UCI tabular datasets. The project is available at https://github.com/DryDew/AutoDV.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 986e66ad-723d-4ee6-b395-4e3a239f8a9eBuilds on14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng et al.ICML 2020 · 1,388 citations
- Recipe for a General, Powerful, Scalable Graph TransformerLadislav Rampásek, Michael Galkin, Vijay Prakash Dwivedi, Anh Tuan Luu et al.NeurIPS 2022 · 1,216 citations
- GraphSAINT: Graph Sampling Based Inductive Learning MethodHanqing Zeng, Hongkuan Zhou, Ajitesh Srivastava, Rajgopal Kannan et al.ICLR 2020 · 1,155 citations
- Graph Inductive Biases in Transformers without Message PassingLiheng Ma, Chen Lin, Derek Lim, Adriana Romero-Soriano et al.ICML 2023 · 185 citations
Related papers
- SpaceMAP: Visualizing High-Dimensional Data by Space ExpansionXinrui Zu, Qian TaoICML 2022 · 12 citations
- Federated t-SNE and UMAP for Distributed Data VisualizationDong Qiao, Xinxian Ma, Jicong FanAAAI 2025 · 3 citations
- Unsupervised visualization of image datasets using contrastive learningJan Niklas Böhm, Philipp Berens, Dmitry KobakICLR 2023 · 6 citations
- Hierarchical Nearest Neighbor Graph Embedding for Efficient Dimensionality ReductionM. Saquib Sarfraz, Marios Koulakis, Constantin Seibold, Rainer StiefelhagenCVPR 2022 · 12 citations
- Joint t-SNE for Comparable Projections of Multiple High-Dimensional DatasetsYinqiao Wang, Lu Chen, Jaemin Jo, Yunhai WangIEEE VIS 2021 · 33 citations
