DBSCOUT: A Density-based Method for Scalable Outlier Detection in Very Large Datasets
Matteo Corain, Paolo Garza, Abolfazl Asudeh
摘要
Recent technological advancements have enabled generating and collecting huge amounts of data in a daily manner. This data is used for different purposes that may impact us on an unprecedented scale. Understanding the data, including detecting its outliers, is a critical step before utilizing it.
Outlier detection has been studied well in the literature but the existing approaches fail to scale to these very large settings. In this paper, we propose DBSCOUT, an efficient exact algorithm for outlier detection with a linear complexity that can run in parallel over multiple independent machines, making it a fit for the settings with billions of tuples. Besides the theoretical analysis, our experiment results confirm orders of magnitude improvement over the existing work, proving the efficiency, scalability, and effectiveness of our approach.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper2
相关 Paper
- Fast and Exact Outlier Detection in Metric Spaces: A Proximity Graph-based ApproachDaichi Amagata, Makoto Onizuka, Takahiro HaraSIGMOD 2021 · 被引用 22 次
- Scalable Algorithms for Densest Subgraph DiscoveryWensheng Luo, Zhuo Tang, Yixiang Fang, Chenhao Ma 等ICDE 2023 · 被引用 14 次
- Dupin: A Parallel Framework for Densest Subgraph Discovery in Fraud Detection on Massive GraphsJiaxin Jiang, Siyuan Yao, Yuchen Li, Qiange Wang 等SIGMOD 2025 · 被引用 3 次
- Fast Density-Peaks Clustering: Multicore-based Parallelization ApproachDaichi Amagata, Takahiro HaraSIGMOD 2021 · 被引用 21 次
- Scalable Grid-based Computation of Kendall's Tau CorrelationNikolaos Koutroumanis, Petros Karampas, Alexandros Karakasidis, Nikos Mamoulis 等VLDB 2026
