Automated, Unsupervised, and Auto-Parameterized Inference of Data Patterns and Anomaly Detection
Qiaolin Qin, Heng Li, Ettore Merlo, Maxime Lamothe
摘要
With the advent of data-centric and machine learning (ML) systems, data quality is playing an increasingly critical role for ensuring the overall quality of software systems. Data preparation, an essential step towards high data quality, is known to be a highly effort-intensive process. Although prior studies have dealt with one of the most impacting issues, data pattern violations, these studies usually require data-specific configurations (i.e., parameterized) or use carefully curated data as learning examples (i.e., supervised), relying on domain knowledge and deep understanding of the data, or demanding significant manual effort. In this paper, we introduce RIOLU: Regex Inferencer autO-parameterized Learning with Uncleaned data. RIOLU is fully automated, automatically parameterized, and does not need labeled samples. RIOLU can generate precise patterns from datasets in various domains, with a high F1 score of 97.2 %, exceeding the state-of-the-art baseline. In addition, according to our experiment on five datasets with anomalies, RIOLU can automatically estimate a data column's error rate, draw normal patterns, and predict anomalies from unlabeled data with higher performance (up to <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"></tex> improvement in terms of F1) than the state-of-the-art baseline, even outperforming ChatGPT in terms of both accuracy (12.3 % higher F1) and efficiency (10 % less inference time). A variant of RIOLU, with user guidance, can further boost its precision, with up to <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"></tex> improvement in terms of F1. Our evaluation in an industrial setting further demonstrates the practical benefits of RIOLU.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- DeepLog: Anomaly Detection and Diagnosis from System Logs through Deep LearningMin Du, Feifei Li, Guineng Zheng, Vivek SrikumarCCS 2017 · 被引用 1,823 次
- Log-based Anomaly Detection with Deep Learning: How Far Are We?Van-Hoang Le, Hongyu ZhangICSE 2022 · 被引用 212 次
- Data Quality for Software Vulnerability DatasetsRoland Croft, Muhammad Ali Babar, M. Mehdi KholoosiICSE 2023 · 被引用 138 次
- Are we building on the rock? on the importance of data preprocessing for code summarizationLin Shi, Fangwen Mu, Xiao Chen, Song Wang 等FSE 2022 · 被引用 65 次
- Mining input grammars from dynamic control flowRahul Gopinath, Björn Mathis, Andreas ZellerFSE 2020 · 被引用 60 次
相关 Paper
- DataVinci: Learning Syntactic and Semantic String RepairsMukul Singh, José Cambronero, Sumit Gulwani, Vu Le 等SIGMOD 2025 · 被引用 3 次
- Auto-Validate: Unsupervised Data Validation Using Data-Domain Patterns Inferred from Data LakesJie Song, Yeye HeSIGMOD 2021 · 被引用 26 次
- Human-in-the-loop Regular Expression Extraction for Single Column Format InconsistencyShaochen Yu, Lei Han, Marta Indulska, Shazia Sadiq 等WWW 2023 · 被引用 3 次
- InfeRE: Step-by-Step Regex Generation via Chain of InferenceShuai Zhang, Xiaodong Gu, Yuting Chen, Beijun ShenASE 2023 · 被引用 8 次
- CtxPipe: Context-aware Data Preparation Pipeline Construction for Machine LearningHaotian Gao, Shaofeng Cai, Tien Tuan Anh Dinh, Zhiyong Huang 等SIGMOD 2025 · 被引用 7 次
