DendroMap: Visual Exploration of Large-Scale Image Datasets for Machine Learning with Treemaps
Donald Bertucci, Md Montaser Hamid, Yashwanthi Anand, Anita Ruangrotsakun, Delyar Tabatabai, Melissa Perez, Minsuk Kahng
摘要
In this paper, we present DendroMap, a novel approach to interactively exploring large-scale image datasets for machine learning (ML). ML practitioners often explore image datasets by generating a grid of images or projecting high-dimensional representations of images into 2-D using dimensionality reduction techniques (e.g., t-SNE). However, neither approach effectively scales to large datasets because images are ineffectively organized and interactions are insufficiently supported. To address these challenges, we develop DendroMap by adapting Treemaps, a well-known visualization technique. DendroMap effectively organizes images by extracting hierarchical cluster structures from high-dimensional representations of images. It enables users to make sense of the overall distributions of datasets and interactively zoom into specific areas of interests at multiple levels of abstraction. Our case studies with widely-used image datasets for deep learning demonstrate that users can discover insights about datasets and trained models by examining the diversity of images, identifying underperforming subgroups, and analyzing classification errors. We conducted a user study that evaluates the effectiveness of DendroMap in grouping and searching tasks by comparing it with a gridified version of t-SNE and found that participants preferred DendroMap. DendroMap is available at https://div-lab.github.io/dendromap/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- PromptMagician: Interactive Prompt Engineering for Text-to-Image CreationYingchaojie Feng, Xingbo Wang, Kamkwai Wong, Sijia Wang 等IEEE VIS 2023 · 被引用 127 次
- Zeno: An Interactive Framework for Behavioral Evaluation of Machine LearningÁngel Alexander Cabrera, Erica Fu, Donald Bertucci, Kenneth Holstein 等CHI 2023 · 被引用 51 次
- A Unified Interactive Model Evaluation for Classification, Object Detection, and Instance Segmentation in Computer VisionChangjian Chen, Yukai Guo, Fengyuan Tian, Shilong Liu 等IEEE VIS 2023 · 被引用 28 次
- Cluster-Aware Grid LayoutYuxing Zhou, Weikai Yang, Jiashu Chen, Changjian Chen 等IEEE VIS 2023 · 被引用 13 次
- Talaria: Interactively Optimizing Machine Learning Models for Efficient InferenceFred Hohman, Chaoqun Wang, Jinmook Lee, Jochen Görtler 等CHI 2024 · 被引用 8 次
它引用的顶会 Paper5
- DECE: Decision Explorer with Counterfactual Explanations for Machine Learning ModelsFurui Cheng, Yao Ming, Huamin QuIEEE VIS 2020 · 被引用 118 次
- Human-in-the-loop Extraction of Interpretable Concepts in Deep Learning ModelsZhenge Zhao, Panpan Xu, Carlos Scheidegger, Liu RenIEEE VIS 2021 · 被引用 59 次
- Where Can We Help? A Visual Analytics Approach to Diagnosing and Improving Semantic Segmentation of Movable ObjectsWenbin He, Lincan Zou, Arvind Kumar Shekar, Liang Gou 等IEEE VIS 2021 · 被引用 45 次
- Crowdsourcing the Perception of Machine TeachingJonggi Hong, Kyungjun Lee, June Xu, Hernisa KacorriCHI 2020 · 被引用 30 次
- II-20: Intelligent and pragmatic analytic categorization of image collectionsJan Zahálka, Marcel Worring, Jarke J. van WijkIEEE VIS 2020 · 被引用 2 次
相关 Paper
- NeuroCartography: Scalable Automatic Visual Summarization of Concepts in Deep Neural NetworksHaekyu Park, Nilaksh Das, Rahul Duggal, Austin P. Wright 等IEEE VIS 2021 · 被引用 28 次
- Revealing the Gap: Visual Comparison of Large-Scale Datasets via Multi-Scale Density Difference MapXinyuan Guo, Xu Zhu, Yilin Ye, Shixia LiuCHI 2026 · 被引用 1 次
- DKMap: Interactive Exploration of Vision-Language Alignment in Multimodal Embeddings via Dynamic Kernel Enhanced ProjectionYilin Ye, Chenxi Ruan, Yu Zhang, Zikun Deng 等IEEE VIS 2025 · 被引用 1 次
- Interactive Visual Study of Multiple Attributes Learning Model of X-Ray Scattering ImagesXinyi Huang, Suphanut Jamonnak, Ye Zhao, Boyu Wang 等IEEE VIS 2020 · 被引用 10 次
- SpaceMAP: Visualizing High-Dimensional Data by Space ExpansionXinrui Zu, Qian TaoICML 2022 · 被引用 12 次
