Human-in-the-loop Outlier Detection
Chengliang Chai, Lei Cao, Guoliang Li, Jian Li, Yuyu Luo, Samuel Madden
Abstract
Outlier detection is critical to a large number of applications from finance fraud detection to health care. Although numerous approaches have been proposed to automatically detect outliers, such outliers detected based on statistical rarity do not necessarily correspond to the true outliers to the interest of applications. In this work, we propose a human-in-the-loop outlier detection approach HOD that effectively leverages human intelligence to discover the true outliers. There are two main challenges in HOD. The first is to design human-friendly questions such that humans can easily understand the questions even if humans know nothing about the outlier detection techniques. The second is to minimize the number of questions. To address the first challenge, we design a clustering-based method to effectively discover a small number of objects that are unlikely to be outliers (aka, inliers) and yet effectively represent the typical characteristics of the given dataset. HOD then leverages this set of inliers (called context inliers) to help humans understand the context in which the outliers occur. This ensures humans are able to easily identify the true outliers from the outlier candidates produced by the machine-based outlier detection techniques. To address the second challenge, we propose a bipartite graph-based question selection strategy that is theoretically proven to be able to minimize the number of questions needed to cover all outlier candidates. Our experimental results on real data sets show that HOD significantly outperforms the state-of-the-art methods on both human efforts and the quality of the discovered outliers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 223196fc-3948-4193-bb09-13d22045b3e7Cited by top-tier papers14
- Synthesizing Natural Language to Visualization (NL2VIS) Benchmarks from NL2SQL BenchmarksYuyu Luo, Nan Tang, Guoliang Li, Chengliang Chai et al.SIGMOD 2021 · 90 citations
- Selective Data Acquisition in the Wild for Model ChargingChengliang Chai, Jiabin Liu, Nan Tang, Guoliang Li et al.VLDB 2022 · 62 citations
- GoodCore: Data-effective and Data-efficient Machine Learning through Coreset Selection over Incomplete DataChengliang Chai, Jiabin Liu, Nan Tang, Ju Fan et al.SIGMOD 2023 · 37 citations
- LakeBench: A Benchmark for Discovering Joinable and Unionable Tables in Data LakesYuhao Deng, Chengliang Chai, Lei Cao, Qin Yuan et al.VLDB 2024 · 36 citations
- Coresets over Multiple Tables for Feature-rich and Data-efficient Machine LearningJiayi Wang, Chengliang Chai, Nan Tang, Jiabin Liu et al.VLDB 2023 · 31 citations
Builds on1
Related papers
- AutoOD: Automatic Outlier DetectionLei Cao, Yizhou Yan, Yu Wang, Samuel Madden et al.SIGMOD 2023 · 9 citations
- ODHD: one-class brain-inspired hyperdimensional computing for outlier detectionRuixuan Wang, Xun Jiao, X. Sharon HuDAC 2022 · 18 citations
- Automatic Unsupervised Outlier Model SelectionYue Zhao, Ryan A. Rossi, Leman AkogluNeurIPS 2021 · 104 citations
- RFOD: Random Forest-Based Outlier Detection for Mixed-Type Tabular DataYihao Ang, Peicheng Yao, Yifan Bao, Yushuo Feng et al.ICDE 2026
- HGOE: Hybrid External and Internal Graph Outlier Exposure for Graph Out-of-Distribution DetectionJunwei He, Qianqian Xu, Yangbangyan Jiang, Zitai Wang et al.ACM MM 2024 · 4 citations
