Large-Scale Evaluation of Topic Models and Dimensionality Reduction Methods for 2D Text Spatialization
Daniel Atzberger, Tim Cech, Matthias Trapp, Rico Richter, Willy Scheibel, Jürgen Döllner, Tobias Schreck
摘要
Topic models are a class of unsupervised learning algorithms for detecting the semantic structure within a text corpus. Together with a subsequent dimensionality reduction algorithm, topic models can be used for deriving spatializations for text corpora as two-dimensional scatter plots, reflecting semantic similarity between the documents and supporting corpus analysis. Although the choice of the topic model, the dimensionality reduction, and their underlying hyperparameters significantly impact the resulting layout, it is unknown which particular combinations result in high-quality layouts with respect to accuracy and perception metrics. To investigate the effectiveness of topic models and dimensionality reduction methods for the spatialization of corpora as two-dimensional scatter plots (or basis for landscape-type visualizations), we present a large-scale, benchmark-based computational evaluation. Our evaluation consists of (1) a set of corpora, (2) a set of layout algorithms that are combinations of topic models and dimensionality reductions, and (3) quality metrics for quantifying the resulting layout. The corpora are given as document-term matrices, and each document is assigned to a thematic class. The chosen metrics quantify the preservation of local and global properties and the perceptual effectiveness of the two-dimensional scatter plots. By evaluating the benchmark on a computing cluster, we derived a multivariate dataset with over 45 000 individual layouts and corresponding quality metrics. Based on the results, we propose guidelines for the effective design of text spatializations that are based on topic models and dimensionality reductions. As a main result, we show that interpretable topic models are beneficial for capturing the structure of text corpora. We furthermore recommend the use of t-SNE as a subsequent dimensionality reduction.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Unveiling High-dimensional Backstage: A Survey for Reliable Visual Analytics with Dimensionality ReductionHyeon Jeon, Hyunwook Lee, Yun-Hsin Kuo, Taehyun Yang 等CHI 2025 · 被引用 29 次
- A Large-Scale Sensitivity Analysis on Latent Embeddings and Dimensionality Reductions for Text SpatializationsDaniel Atzberger, Tim Cech, Willy Scheibel, Jürgen Döllner 等IEEE VIS 2024 · 被引用 8 次
它引用的顶会 Paper4
- Revisiting Dimensionality Reduction Techniques for Visual Cluster Analysis: An Empirical StudyJiazhi Xia, Yuchen Zhang, Jie Song, Yang Chen 等IEEE VIS 2021 · 被引用 82 次
- Interactive Dimensionality Reduction for Comparative AnalysisTakanori Fujiwara, Xinhai Wei, Jian Zhao, Kwan-Liu MaIEEE VIS 2021 · 被引用 47 次
- Interactive Visual Cluster Analysis by Contrastive Dimensionality ReductionJiazhi Xia, Linquan Huang, Weixing Lin, Xin Zhao 等IEEE VIS 2022 · 被引用 38 次
- Predicting User Preferences of Dimensionality Reduction Embedding QualityCristina Morariu, Adrien Bibal, René Cutura, Benoît Frénay 等IEEE VIS 2022 · 被引用 13 次
相关 Paper
- Uncovering How Scatterplot Features Skew Visual Class SeparationS. Sandra Bae, Takanori Fujiwara, Chin Tseng, Danielle Albers SzafirCHI 2025 · 被引用 3 次
- Evaluation of Thematic Coherence in MicroblogsIman Munire Bilal, Bo Wang, Maria Liakata, Rob Procter 等ACL 2021
- Representing Mixtures of Word Embeddings with Mixtures of Topic EmbeddingsDongsheng Wang, Dandan Guo, He Zhao, Huangjie Zheng 等ICLR 2022 · 被引用 56 次
- Large-Scale Correlation Analysis of Automated Metrics for Topic ModelsJia Peng Lim, Hady W. LauwACL 2023 · 被引用 13 次
- User Ex Machina : Simulation as a Design Probe in Human-in-the-Loop Text AnalyticsAnamaria Crisan, Michael CorrellCHI 2021 · 被引用 11 次
