Divisi: Interactive Search and Visualization for Scalable Exploratory Subgroup Analysis
Venkatesh Sivaraman, Zexuan Li, Adam Perer
Abstract
Re-rank subgroups by criteria such as rate and coverage of an outcome B Visualize subgroup overlap in a map of the dataset E Selected subgroups are designated with consistent colors throughout interface Ranking functions and metrics can be defined in Python code or interactively Interface embedded in a computational notebook for frictionless setup Compare subgroup metrics at a glance
C Edit subgroup definitions to see how alternative values affect metrics D Run algorithm to find data subgroups with interesting differences A Save important subgroups for further review F Figure 1: Divisi is an interactive visualization system to help data scientists perform exploratory subgroup analysis on large datasets with many feature dimensions, such as the dataset of airline passenger satisfaction ratings shown [1]. Implemented as a computational notebook widget, Divisi includes a novel approximate subgroup discovery algorithm (A) which allows interactive re-ranking by customizable functions, such as error rate and coverage (B). Users can compare metrics across subgroups (C) and test alternative rule definitions (D) to evaluate subgroups. Finally, the Subgroup Map (E) depicts overlap and coverage between groups, so users can curate the most representative subgroups for review (F).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Tempo: Helping Data Scientists and Domain Experts Collaboratively Specify Predictive Modeling TasksVenkatesh Sivaraman, Anika Vaishampayan, Xiaotong Li, Brian R. Buck et al.CHI 2025 · 1 citation
- BloomIntent: Automating Search Evaluation with LLM-Generated Fine-Grained User IntentsYoonseo Choi, Eunhye Kim, Hyunwoo Kim, Donghyun Park et al.UIST 2025 · 1 citation
- Evalet: Evaluating Large Language Models through Functional FragmentationTae Soo Kim, Heechan Lee, Yoonjoo Lee, Joseph Seering et al.CHI 2026 · 1 citation
- Data-Prompt Co-Evolution: Growing Test Sets to Refine LLM BehaviorMinjae Lee, Minsuk KahngCHI 2026 · 1 citation
Builds on17
- Domino: Discovering Systematic Errors with Cross-Modal EmbeddingsSabri Eyuboglu, Maya Varma, Khaled Kamal Saab, Jean-Benoit Delbrouck et al.ICLR 2022 · 178 citations
- Lux: Always-on Visualization Recommendations for Exploratory Dataframe WorkflowsDoris Jung Lin Lee, Dixin Tang, Kunal Agarwal, Thyne Boonmark et al.VLDB 2022 · 61 citations
- Discovering and Validating AI Errors With Crowdsourced Failure ReportsÁngel Alexander Cabrera, Abraham J. Druck, Jason I. Hong, Adam PererCSCW 2021 · 60 citations
- Zeno: An Interactive Framework for Behavioral Evaluation of Machine LearningÁngel Alexander Cabrera, Erica Fu, Donald Bertucci, Kenneth Holstein et al.CHI 2023 · 51 citations
- Looking for Trouble: Analyzing Classifier Behavior via Pattern DivergenceEliana Pastor, Luca de Alfaro, Elena BaralisSIGMOD 2021 · 51 citations
Related papers
- DIVI: Dynamically Interactive VisualizationLuke S. Snyder, Jeffrey HeerIEEE VIS 2023 · 14 citations
- VLSlice: Interactive Vision-and-Language Slice DiscoveryEric Slyman, Minsuk Kahng, Stefan LeeICCV 2023 · 11 citations
- Query-Driven Data Exploration with Heterogeneous Treatment EffectsAntonis Mandamadiotis, Sihem Amer-Yahia, Georgia KoutrikaICDE 2026
- Robust Recursive Partitioning for Heterogeneous Treatment Effects with Uncertainty QuantificationHyun-Suk Lee, Yao Zhang, William R. Zame, Cong Shen et al.NeurIPS 2020 · 21 citations
- DataPilot: Utilizing Quality and Usage Information for Subset Selection during Visual Data PreparationArpit Narechania, Fan Du, Atanu R. Sinha, Ryan A. Rossi et al.CHI 2023 · 13 citations
