Tailoring Data Source Distributions for Fairness-aware Data Integration
Fatemeh Nargesian, Abolfazl Asudeh, H. V. Jagadish
Abstract
Data scientists often develop data sets for analysis by drawing upon sources of data available to them. A major challenge is to ensure that the data set used for analysis has an appropriate representation of relevant (demographic) groups: it meets desired distribution requirements. Whether data is collected through some experiment or obtained from some data provider, the data from any single source may not meet the desired distribution requirements. Therefore, a union of data from multiple sources is often required. In this paper, we study how to acquire such data in the most cost effective manner, for typical cost functions observed in practice. We present an optimal solution for binary groups when the underlying distributions of data sources are known and all data sources have equal costs. For the generic case with unequal costs, we design an approximation algorithm that performs well in practice. When the underlying distributions are unknown, we develop an exploration-exploitation based strategy with a reward function that captures the cost and approximations of group distributions in each data source. Besides theoretical analysis, we conduct comprehensive experiments that confirm the effectiveness of our algorithms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e425f9b4-96fd-45f4-ab7c-907288d1c172Cited by top-tier papers15
- Selective Data Acquisition in the Wild for Model ChargingChengliang Chai, Jiabin Liu, Nan Tang, Guoliang Li et al.VLDB 2022 · 62 citations
- Through the Fairness Lens: Experimental Analysis and Evaluation of Entity MatchingNima Shahbazi, Nikola Danevski, Fatemeh Nargesian, Abolfazl Asudeh et al.VLDB 2023 · 21 citations
- Maximizing Fair Content Spread via Edge Suggestion in Social NetworksIan P. Swift, Sana Ebrahimi, Azade Nova, Abolfazl AsudehVLDB 2022 · 19 citations
- Fairness-Aware Range Queries for Selecting Unbiased DataSuraj Shetiya, Ian P. Swift, Abolfazl Asudeh, Gautam DasICDE 2022 · 19 citations
- Towards Distribution-aware Query Answering in Data MarketsAbolfazl Asudeh, Fatemeh NargesianVLDB 2022 · 17 citations
Builds on4
- Rank Aggregation Algorithms for Fair ConsensusCaitlin Kuhlman, Elke A. RundensteinerVLDB 2020 · 60 citations
- Identifying Insufficient Data Coverage in Databases with Multiple RelationsYin Lin, Yifan Guan, Abolfazl Asudeh, H. V. JagadishVLDB 2020 · 50 citations
- Identifying Insufficient Data Coverage for Ordinal Continuous-Valued AttributesAbolfazl Asudeh, Nima Shahbazi, Zhongjun Jin, H. V. JagadishSIGMOD 2021 · 30 citations
- ARDA: Automatic Relational Data Augmentation for Machine LearningNadiia Chepurko, Ryan Marcus, Emanuel Zgraggen, Raul Castro Fernandez et al.VLDB 2020 · 14 citations
Related papers
- One Size Does Not Fit All: A Bandit-Based Sampler Combination Framework with Theoretical GuaranteesJinglin Peng, Bolin Ding, Jiannan Wang, Kai Zeng et al.SIGMOD 2022 · 7 citations
- Approximating a Distribution Using Weight QueriesNadav Barak, Sivan SabatoICML 2021
- CDA: Cost-Sensitive Data Acquisition for Incomplete DatasetsKaiyu Li, Xiaohui Yu, Jian PeiICDE 2025
- Intersectional Fairness in Reinforcement Learning with Large State and Constraint SpacesEric Eaton, Marcel Hussing, Michael Kearns, Aaron Roth et al.ICML 2025
- Worst-Case Analysis for Randomly Collected DataJustin Y. Chen, Gregory Valiant, Paul ValiantNeurIPS 2020 · 4 citations
