Outlier Summarization via Human Interpretable Rules
Yuhao Deng, Yu Wang, Lei Cao, Lianpeng Qiao, Yuping Wang, Xu Jingzhe, Yizhou Yan, Samuel Madden
摘要
Outlier detection is crucial for preventing financial fraud, network intrusions, and device failures. Users often expect systems to automatically summarize and interpret outlier detection results to reduce human effort and convert outliers into actionable insights. However, existing methods fail to effectively assist users in identifying the root causes of outliers, as they only pinpoint data attributes without considering outliers in the same subspace may have different causes.
To fill this gap, we propose STAIR, which learns concise and human-understandable rules to summarize and explain outlier detection results with finer granularity. These rules consider both attributes and associated values. STAIR employs an interpretation-aware optimization objective to generate a small number of rules with minimal complexity for strong interpretability. The learning algorithm of STAIR produces a rule set by iteratively splitting the large rules and is optimal in maximizing this objective in each iteration. Moreover, to effectively handle high dimensional, highly complex data sets that are hard to summarize with simple rules, we propose a localized STAIR approach, called L-STAIR. Taking data locality into consideration, it simultaneously partitions data and learns a set of localized rules for each partition. Our experimental study on many outlier benchmark datasets shows that STAIR significantly reduces the complexity of the rules required to summarize the outlier detection results, thus more amenable for humans to understand and evaluate.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- GraphMaster: Automated Graph Synthesis via LLM Agents in Data-Limited EnvironmentsEnjun Du, Xunkai Li, Tian Jin, Zhihan Zhang 等NeurIPS 2025 · 被引用 25 次
- Finding Non-Redundant Simpson's Paradox in Multidimensional DataYi Yang, Jian Pei, Jun Yang, Jichun XieVLDB 2026
它引用的顶会 Paper4
- Human-in-the-loop Outlier DetectionChengliang Chai, Lei Cao, Guoliang Li, Jian Li 等SIGMOD 2020 · 被引用 57 次
- Efficient and Effective Data Imputation with Influence FunctionsXiaoye Miao, Yangyang Wu, Lu Chen, Yunjun Gao 等VLDB 2022 · 被引用 38 次
- GoodCore: Data-effective and Data-efficient Machine Learning through Coreset Selection over Incomplete DataChengliang Chai, Jiabin Liu, Nan Tang, Ju Fan 等SIGMOD 2023 · 被引用 37 次
- MisDetect: Iterative Mislabel Detection using Early LossYuhao Deng, Chengliang Chai, Lei Cao, Nan Tang 等VLDB 2024 · 被引用 13 次
相关 Paper
- Beyond Outlier Detection: Outlier Interpretation by Attention-Guided Triplet Deviation NetworkHongzuo Xu, Yijie Wang, Songlei Jian, Zhenyu Huang 等WWW 2021 · 被引用 40 次
- Robust and Explainable Autoencoders for Unsupervised Time Series Outlier DetectionTung Kieu, Bin Yang, Chenjuan Guo, Christian S. Jensen 等ICDE 2022 · 被引用 60 次
- Interpreting Unsupervised Anomaly Detection in Security via Rule ExtractionRuoyu Li, Qing Li, Yu Zhang, Dan Zhao 等NeurIPS 2023 · 被引用 18 次
- xNIDS: Explaining Deep Learning-based Network Intrusion Detection Systems for Active Intrusion ResponsesFeng Wei, Hongda Li, Ziming Zhao, Hongxin HuUSENIX Security 2023
- Understanding Failures of Deep Networks via Robust Feature ExtractionSahil Singla, Besmira Nushi, Shital Shah, Ece Kamar 等CVPR 2021
