Can LLMs Bridge Domain and Visualization? A Case Study on High-Dimension Data Visualization in Single-Cell Transcriptomics
Qianwen Wang, Xinyi Liu, Nils Gehlenborg
摘要
While many visualizations are built for domain users (e.g., biologists, machine learning developers), understanding how visualizations are used in the domain has long been a challenging task. Previous research has relied on either interviewing a limited number of domain users or reviewing relevant application papers in the visualization community, neither of which provides comprehensive insight into visualizations "in the wild" of a specific domain. This paper aims to fill this gap by examining the potential of using Large Language Models (LLM) to analyze visualization usage in domain literature. We use high-dimension (HD) data visualization in sing-cell transcriptomics as a test case, analyzing 1,203 papers that describe 2,056 HD visualizations with highly specialized domain terminologies (e.g., biomarkers, cell lineage). To facilitate this analysis, we introduce a human-in-the-loop LLM workflow that can effectively analyze a large collection of papers and translate domain-specific terminology into standardized data and task abstractions. Instead of relying solely on LLMs for end-to-end analysis, our workflow enhances analytical quality through 1) integrating image processing and traditional NLP methods to prepare well-structured inputs for three targeted LLM subtasks (i.e., translating domain terminology, summarizing analysis tasks, and performing categorization), and 2) establishing checkpoints for human involvement and validation throughout the process. The analysis results, validated with expert interviews and a test set, revealed three often overlooked aspects in HD visualization: trajectories in HD spaces, inter-cluster relationships, and dimension clustering. This research provides a stepping stone for future studies seeking to use LLMs to bridge the gap between visualization design and domain-specific usage.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Cell2Sentence: Teaching Large Language Models the Language of BiologyDaniel LeVine, Syed Asad Rizvi, Sacha Lévy, Nazreen Pallikkavaliyaveetil 等ICML 2024 · 被引用 70 次
- LLM4Cell: Taxonomy and Evaluation of LLM and Agentic Models for Single-Cell BiologySajib Acharjee Dip, Adrika Zafor, Bikash Kumar Paul, Uddip Acharjee Shuvo 等ACL 2026
- VisEval: A Benchmark for Data Visualization in the Era of Large Language ModelsNan Chen, Yuge Zhang, Jiahang Xu, Kan Ren 等IEEE VIS 2024 · 被引用 44 次
- ClimateSOM: A Visual Analysis Workflow for Climate Ensemble DatasetsYuya Kawakami, Daniel Cayan, Dongyu Liu, Kwan-Liu MaIEEE VIS 2025 · 被引用 2 次
- Text2Vis: A Challenging and Diverse Benchmark for Generating Multimodal Visualizations from TextMizanur Rahman, Md. Tahmid Rahman Laskar, Shafiq Joty, Enamul HoqueEMNLP 2025 · 被引用 1 次
