Outlier Summarization via Human Interpretable Rules
Yuhao Deng, Yu Wang, Lei Cao, Lianpeng Qiao, Yuping Wang, Xu Jingzhe, Yizhou Yan, Samuel Madden
Abstract
Outlier detection is crucial for preventing financial fraud, network intrusions, and device failures. Users often expect systems to automatically summarize and interpret outlier detection results to reduce human effort and convert outliers into actionable insights. However, existing methods fail to effectively assist users in identifying the root causes of outliers, as they only pinpoint data attributes without considering outliers in the same subspace may have different causes.
To fill this gap, we propose STAIR, which learns concise and human-understandable rules to summarize and explain outlier detection results with finer granularity. These rules consider both attributes and associated values. STAIR employs an interpretation-aware optimization objective to generate a small number of rules with minimal complexity for strong interpretability. The learning algorithm of STAIR produces a rule set by iteratively splitting the large rules and is optimal in maximizing this objective in each iteration. Moreover, to effectively handle high dimensional, highly complex data sets that are hard to summarize with simple rules, we propose a localized STAIR approach, called L-STAIR. Taking data locality into consideration, it simultaneously partitions data and learns a set of localized rules for each partition. Our experimental study on many outlier benchmark datasets shows that STAIR significantly reduces the complexity of the rules required to summarize the outlier detection results, thus more amenable for humans to understand and evaluate.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8e8eaa83-25e8-48b1-bfe5-27f3cfa0f813Cited by top-tier papers2
- GraphMaster: Automated Graph Synthesis via LLM Agents in Data-Limited EnvironmentsEnjun Du, Xunkai Li, Tian Jin, Zhihan Zhang et al.NeurIPS 2025 · 25 citations
- Finding Non-Redundant Simpson's Paradox in Multidimensional DataYi Yang, Jian Pei, Jun Yang, Jichun XieVLDB 2026
Builds on4
- Human-in-the-loop Outlier DetectionChengliang Chai, Lei Cao, Guoliang Li, Jian Li et al.SIGMOD 2020 · 57 citations
- Efficient and Effective Data Imputation with Influence FunctionsXiaoye Miao, Yangyang Wu, Lu Chen, Yunjun Gao et al.VLDB 2022 · 38 citations
- GoodCore: Data-effective and Data-efficient Machine Learning through Coreset Selection over Incomplete DataChengliang Chai, Jiabin Liu, Nan Tang, Ju Fan et al.SIGMOD 2023 · 37 citations
- MisDetect: Iterative Mislabel Detection using Early LossYuhao Deng, Chengliang Chai, Lei Cao, Nan Tang et al.VLDB 2024 · 13 citations
Related papers
- Beyond Outlier Detection: Outlier Interpretation by Attention-Guided Triplet Deviation NetworkHongzuo Xu, Yijie Wang, Songlei Jian, Zhenyu Huang et al.WWW 2021 · 40 citations
- Robust and Explainable Autoencoders for Unsupervised Time Series Outlier DetectionTung Kieu, Bin Yang, Chenjuan Guo, Christian S. Jensen et al.ICDE 2022 · 60 citations
- Interpreting Unsupervised Anomaly Detection in Security via Rule ExtractionRuoyu Li, Qing Li, Yu Zhang, Dan Zhao et al.NeurIPS 2023 · 18 citations
- xNIDS: Explaining Deep Learning-based Network Intrusion Detection Systems for Active Intrusion ResponsesFeng Wei, Hongda Li, Ziming Zhao, Hongxin HuUSENIX Security 2023
- Understanding Failures of Deep Networks via Robust Feature ExtractionSahil Singla, Besmira Nushi, Shital Shah, Ece Kamar et al.CVPR 2021
