Divisi: Interactive Search and Visualization for Scalable Exploratory Subgroup Analysis
Venkatesh Sivaraman, Zexuan Li, Adam Perer
摘要
Re-rank subgroups by criteria such as rate and coverage of an outcome B Visualize subgroup overlap in a map of the dataset E Selected subgroups are designated with consistent colors throughout interface Ranking functions and metrics can be defined in Python code or interactively Interface embedded in a computational notebook for frictionless setup Compare subgroup metrics at a glance
C Edit subgroup definitions to see how alternative values affect metrics D Run algorithm to find data subgroups with interesting differences A Save important subgroups for further review F Figure 1: Divisi is an interactive visualization system to help data scientists perform exploratory subgroup analysis on large datasets with many feature dimensions, such as the dataset of airline passenger satisfaction ratings shown [1]. Implemented as a computational notebook widget, Divisi includes a novel approximate subgroup discovery algorithm (A) which allows interactive re-ranking by customizable functions, such as error rate and coverage (B). Users can compare metrics across subgroups (C) and test alternative rule definitions (D) to evaluate subgroups. Finally, the Subgroup Map (E) depicts overlap and coverage between groups, so users can curate the most representative subgroups for review (F).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Tempo: Helping Data Scientists and Domain Experts Collaboratively Specify Predictive Modeling TasksVenkatesh Sivaraman, Anika Vaishampayan, Xiaotong Li, Brian R. Buck 等CHI 2025 · 被引用 1 次
- BloomIntent: Automating Search Evaluation with LLM-Generated Fine-Grained User IntentsYoonseo Choi, Eunhye Kim, Hyunwoo Kim, Donghyun Park 等UIST 2025 · 被引用 1 次
- Evalet: Evaluating Large Language Models through Functional FragmentationTae Soo Kim, Heechan Lee, Yoonjoo Lee, Joseph Seering 等CHI 2026 · 被引用 1 次
- Data-Prompt Co-Evolution: Growing Test Sets to Refine LLM BehaviorMinjae Lee, Minsuk KahngCHI 2026 · 被引用 1 次
它引用的顶会 Paper17
- Domino: Discovering Systematic Errors with Cross-Modal EmbeddingsSabri Eyuboglu, Maya Varma, Khaled Kamal Saab, Jean-Benoit Delbrouck 等ICLR 2022 · 被引用 178 次
- Lux: Always-on Visualization Recommendations for Exploratory Dataframe WorkflowsDoris Jung Lin Lee, Dixin Tang, Kunal Agarwal, Thyne Boonmark 等VLDB 2022 · 被引用 61 次
- Discovering and Validating AI Errors With Crowdsourced Failure ReportsÁngel Alexander Cabrera, Abraham J. Druck, Jason I. Hong, Adam PererCSCW 2021 · 被引用 60 次
- Zeno: An Interactive Framework for Behavioral Evaluation of Machine LearningÁngel Alexander Cabrera, Erica Fu, Donald Bertucci, Kenneth Holstein 等CHI 2023 · 被引用 51 次
- Looking for Trouble: Analyzing Classifier Behavior via Pattern DivergenceEliana Pastor, Luca de Alfaro, Elena BaralisSIGMOD 2021 · 被引用 51 次
相关 Paper
- DIVI: Dynamically Interactive VisualizationLuke S. Snyder, Jeffrey HeerIEEE VIS 2023 · 被引用 14 次
- VLSlice: Interactive Vision-and-Language Slice DiscoveryEric Slyman, Minsuk Kahng, Stefan LeeICCV 2023 · 被引用 11 次
- Query-Driven Data Exploration with Heterogeneous Treatment EffectsAntonis Mandamadiotis, Sihem Amer-Yahia, Georgia KoutrikaICDE 2026
- Robust Recursive Partitioning for Heterogeneous Treatment Effects with Uncertainty QuantificationHyun-Suk Lee, Yao Zhang, William R. Zame, Cong Shen 等NeurIPS 2020 · 被引用 21 次
- DataPilot: Utilizing Quality and Usage Information for Subset Selection during Visual Data PreparationArpit Narechania, Fan Du, Atanu R. Sinha, Ryan A. Rossi 等CHI 2023 · 被引用 13 次
