DendroMap: Visual Exploration of Large-Scale Image Datasets for Machine Learning with Treemaps
Donald Bertucci, Md Montaser Hamid, Yashwanthi Anand, Anita Ruangrotsakun, Delyar Tabatabai, Melissa Perez, Minsuk Kahng
Abstract
In this paper, we present DendroMap, a novel approach to interactively exploring large-scale image datasets for machine learning (ML). ML practitioners often explore image datasets by generating a grid of images or projecting high-dimensional representations of images into 2-D using dimensionality reduction techniques (e.g., t-SNE). However, neither approach effectively scales to large datasets because images are ineffectively organized and interactions are insufficiently supported. To address these challenges, we develop DendroMap by adapting Treemaps, a well-known visualization technique. DendroMap effectively organizes images by extracting hierarchical cluster structures from high-dimensional representations of images. It enables users to make sense of the overall distributions of datasets and interactively zoom into specific areas of interests at multiple levels of abstraction. Our case studies with widely-used image datasets for deep learning demonstrate that users can discover insights about datasets and trained models by examining the diversity of images, identifying underperforming subgroups, and analyzing classification errors. We conducted a user study that evaluates the effectiveness of DendroMap in grouping and searching tasks by comparing it with a gridified version of t-SNE and found that participants preferred DendroMap. DendroMap is available at https://div-lab.github.io/dendromap/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 197ccc2a-0be5-49a5-8fd0-90fbd2f1dd5cCited by top-tier papers7
- PromptMagician: Interactive Prompt Engineering for Text-to-Image CreationYingchaojie Feng, Xingbo Wang, Kamkwai Wong, Sijia Wang et al.IEEE VIS 2023 · 127 citations
- Zeno: An Interactive Framework for Behavioral Evaluation of Machine LearningÁngel Alexander Cabrera, Erica Fu, Donald Bertucci, Kenneth Holstein et al.CHI 2023 · 51 citations
- A Unified Interactive Model Evaluation for Classification, Object Detection, and Instance Segmentation in Computer VisionChangjian Chen, Yukai Guo, Fengyuan Tian, Shilong Liu et al.IEEE VIS 2023 · 28 citations
- Cluster-Aware Grid LayoutYuxing Zhou, Weikai Yang, Jiashu Chen, Changjian Chen et al.IEEE VIS 2023 · 13 citations
- Talaria: Interactively Optimizing Machine Learning Models for Efficient InferenceFred Hohman, Chaoqun Wang, Jinmook Lee, Jochen Görtler et al.CHI 2024 · 8 citations
Builds on5
- DECE: Decision Explorer with Counterfactual Explanations for Machine Learning ModelsFurui Cheng, Yao Ming, Huamin QuIEEE VIS 2020 · 118 citations
- Human-in-the-loop Extraction of Interpretable Concepts in Deep Learning ModelsZhenge Zhao, Panpan Xu, Carlos Scheidegger, Liu RenIEEE VIS 2021 · 59 citations
- Where Can We Help? A Visual Analytics Approach to Diagnosing and Improving Semantic Segmentation of Movable ObjectsWenbin He, Lincan Zou, Arvind Kumar Shekar, Liang Gou et al.IEEE VIS 2021 · 45 citations
- Crowdsourcing the Perception of Machine TeachingJonggi Hong, Kyungjun Lee, June Xu, Hernisa KacorriCHI 2020 · 30 citations
- II-20: Intelligent and pragmatic analytic categorization of image collectionsJan Zahálka, Marcel Worring, Jarke J. van WijkIEEE VIS 2020 · 2 citations
Related papers
- NeuroCartography: Scalable Automatic Visual Summarization of Concepts in Deep Neural NetworksHaekyu Park, Nilaksh Das, Rahul Duggal, Austin P. Wright et al.IEEE VIS 2021 · 28 citations
- Revealing the Gap: Visual Comparison of Large-Scale Datasets via Multi-Scale Density Difference MapXinyuan Guo, Xu Zhu, Yilin Ye, Shixia LiuCHI 2026 · 1 citation
- DKMap: Interactive Exploration of Vision-Language Alignment in Multimodal Embeddings via Dynamic Kernel Enhanced ProjectionYilin Ye, Chenxi Ruan, Yu Zhang, Zikun Deng et al.IEEE VIS 2025 · 1 citation
- Interactive Visual Study of Multiple Attributes Learning Model of X-Ray Scattering ImagesXinyi Huang, Suphanut Jamonnak, Ye Zhao, Boyu Wang et al.IEEE VIS 2020 · 10 citations
- SpaceMAP: Visualizing High-Dimensional Data by Space ExpansionXinrui Zu, Qian TaoICML 2022 · 12 citations
