Improving Data Imputation Through a Tuned Strategy for Dependency Discovery
Bernardo Breve, Loredana Caruccio, Tullio Pizzuti, Giuseppe Polese
摘要
Data Imputation approaches aim at improving the quality of data trying to infer values often missing into data. Among others, dependency-based imputation approaches exploit data relationships among attributes, such as Relaxed Functional Dependencies relaxing on the attribute comparison , to provide semantically coherent and less biased imputations by exploiting similar candidate tuples. However, according to the possibility of considering variable similarity threshold combinations, existing RFD discovery algorithms can limit imputation quality or become impractical for big datasets. In this paper, we propose triard, a tuned strategy for dependency discovery that iteratively adjusts similarity thresholds based on imputation performance, thus selecting the most effective for selecting tuple candidates to impute missing values. Experimental results on 20 real-world datasets show the improvement in both imputation performance and execution time with respect to the discovered by the algorithm domino. Moreover, triard resulted in a higher imputation performance, mainly in terms of precision, when compared with other Data Imputation approaches.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Efficient Relaxed Functional Dependency Discovery with Minimal Set CoverXiaoou Ding, Yida Liu, Hongzhi Wang, Chen Wang 等ICDE 2024 · 被引用 7 次
- TSDDISCOVER: Discovering Data Dependency for Time Series DataXiaoou Ding, Yingze Li, Hongzhi Wang, Chen Wang 等ICDE 2024 · 被引用 11 次
- Imputing Various Incomplete Attributes via Distance Likelihood MaximizationShaoxu Song, Yu SunKDD 2020 · 被引用 15 次
- Efficient Discovery of Relaxed Functional DependenciesMengran Li, Zijing Tan, Honghui Yang, Shuai MaVLDB 2025
- IndiBits: Incremental Discovery of Relaxed Functional Dependencies using Bitwise SimilarityBernardo Breve, Loredana Caruccio, Stefano Cirillo, Vincenzo Deufemia 等ICDE 2023 · 被引用 7 次
