Lune

ICDE2026Top-tier venue

Improving Data Imputation Through a Tuned Strategy for Dependency Discovery

Bernardo Breve, Loredana Caruccio, Tullio Pizzuti, Giuseppe Polese

2026Year

Abstract

Data Imputation approaches aim at improving the quality of data trying to infer values often missing into data. Among others, dependency-based imputation approaches exploit data relationships among attributes, such as Relaxed Functional Dependencies relaxing on the attribute comparison (RFDcs)\left(\text{RFD}_{c} \mathrm{s}\right), to provide semantically coherent and less biased imputations by exploiting similar candidate tuples. However, according to the possibility of considering variable similarity threshold combinations, existing RFD c_{c} discovery algorithms can limit imputation quality or become impractical for big datasets. In this paper, we propose triard, a tuned strategy for dependency discovery that iteratively adjusts similarity thresholds based on imputation performance, thus selecting the most effective RFDcs\mathbf{R F D}_{c} \mathbf{s} for selecting tuple candidates to impute missing values. Experimental results on 20 real-world datasets show the improvement in both imputation performance and execution time with respect to the RFDcs\mathbf{R F D}_{c} \mathbf{s} discovered by the algorithm domino. Moreover, triard resulted in a higher imputation performance, mainly in terms of precision, when compared with other Data Imputation approaches.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines