Human-in-the-loop Outlier Detection
Chengliang Chai, Lei Cao, Guoliang Li, Jian Li, Yuyu Luo, Samuel Madden
摘要
Outlier detection is critical to a large number of applications from finance fraud detection to health care. Although numerous approaches have been proposed to automatically detect outliers, such outliers detected based on statistical rarity do not necessarily correspond to the true outliers to the interest of applications. In this work, we propose a human-in-the-loop outlier detection approach HOD that effectively leverages human intelligence to discover the true outliers. There are two main challenges in HOD. The first is to design human-friendly questions such that humans can easily understand the questions even if humans know nothing about the outlier detection techniques. The second is to minimize the number of questions. To address the first challenge, we design a clustering-based method to effectively discover a small number of objects that are unlikely to be outliers (aka, inliers) and yet effectively represent the typical characteristics of the given dataset. HOD then leverages this set of inliers (called context inliers) to help humans understand the context in which the outliers occur. This ensures humans are able to easily identify the true outliers from the outlier candidates produced by the machine-based outlier detection techniques. To address the second challenge, we propose a bipartite graph-based question selection strategy that is theoretically proven to be able to minimize the number of questions needed to cover all outlier candidates. Our experimental results on real data sets show that HOD significantly outperforms the state-of-the-art methods on both human efforts and the quality of the discovered outliers.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Synthesizing Natural Language to Visualization (NL2VIS) Benchmarks from NL2SQL BenchmarksYuyu Luo, Nan Tang, Guoliang Li, Chengliang Chai 等SIGMOD 2021 · 被引用 90 次
- Selective Data Acquisition in the Wild for Model ChargingChengliang Chai, Jiabin Liu, Nan Tang, Guoliang Li 等VLDB 2022 · 被引用 62 次
- GoodCore: Data-effective and Data-efficient Machine Learning through Coreset Selection over Incomplete DataChengliang Chai, Jiabin Liu, Nan Tang, Ju Fan 等SIGMOD 2023 · 被引用 37 次
- LakeBench: A Benchmark for Discovering Joinable and Unionable Tables in Data LakesYuhao Deng, Chengliang Chai, Lei Cao, Qin Yuan 等VLDB 2024 · 被引用 36 次
- Coresets over Multiple Tables for Feature-rich and Data-efficient Machine LearningJiayi Wang, Chengliang Chai, Nan Tang, Jiabin Liu 等VLDB 2023 · 被引用 31 次
它引用的顶会 Paper1
相关 Paper
- AutoOD: Automatic Outlier DetectionLei Cao, Yizhou Yan, Yu Wang, Samuel Madden 等SIGMOD 2023 · 被引用 9 次
- ODHD: one-class brain-inspired hyperdimensional computing for outlier detectionRuixuan Wang, Xun Jiao, X. Sharon HuDAC 2022 · 被引用 18 次
- Automatic Unsupervised Outlier Model SelectionYue Zhao, Ryan A. Rossi, Leman AkogluNeurIPS 2021 · 被引用 104 次
- RFOD: Random Forest-Based Outlier Detection for Mixed-Type Tabular DataYihao Ang, Peicheng Yao, Yifan Bao, Yushuo Feng 等ICDE 2026
- HGOE: Hybrid External and Internal Graph Outlier Exposure for Graph Out-of-Distribution DetectionJunwei He, Qianqian Xu, Yangbangyan Jiang, Zitai Wang 等ACM MM 2024 · 被引用 4 次
