Tailoring Data Source Distributions for Fairness-aware Data Integration
Fatemeh Nargesian, Abolfazl Asudeh, H. V. Jagadish
摘要
Data scientists often develop data sets for analysis by drawing upon sources of data available to them. A major challenge is to ensure that the data set used for analysis has an appropriate representation of relevant (demographic) groups: it meets desired distribution requirements. Whether data is collected through some experiment or obtained from some data provider, the data from any single source may not meet the desired distribution requirements. Therefore, a union of data from multiple sources is often required. In this paper, we study how to acquire such data in the most cost effective manner, for typical cost functions observed in practice. We present an optimal solution for binary groups when the underlying distributions of data sources are known and all data sources have equal costs. For the generic case with unequal costs, we design an approximation algorithm that performs well in practice. When the underlying distributions are unknown, we develop an exploration-exploitation based strategy with a reward function that captures the cost and approximations of group distributions in each data source. Besides theoretical analysis, we conduct comprehensive experiments that confirm the effectiveness of our algorithms.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Selective Data Acquisition in the Wild for Model ChargingChengliang Chai, Jiabin Liu, Nan Tang, Guoliang Li 等VLDB 2022 · 被引用 62 次
- Through the Fairness Lens: Experimental Analysis and Evaluation of Entity MatchingNima Shahbazi, Nikola Danevski, Fatemeh Nargesian, Abolfazl Asudeh 等VLDB 2023 · 被引用 21 次
- Maximizing Fair Content Spread via Edge Suggestion in Social NetworksIan P. Swift, Sana Ebrahimi, Azade Nova, Abolfazl AsudehVLDB 2022 · 被引用 19 次
- Fairness-Aware Range Queries for Selecting Unbiased DataSuraj Shetiya, Ian P. Swift, Abolfazl Asudeh, Gautam DasICDE 2022 · 被引用 19 次
- Towards Distribution-aware Query Answering in Data MarketsAbolfazl Asudeh, Fatemeh NargesianVLDB 2022 · 被引用 17 次
它引用的顶会 Paper4
- Rank Aggregation Algorithms for Fair ConsensusCaitlin Kuhlman, Elke A. RundensteinerVLDB 2020 · 被引用 60 次
- Identifying Insufficient Data Coverage in Databases with Multiple RelationsYin Lin, Yifan Guan, Abolfazl Asudeh, H. V. JagadishVLDB 2020 · 被引用 50 次
- Identifying Insufficient Data Coverage for Ordinal Continuous-Valued AttributesAbolfazl Asudeh, Nima Shahbazi, Zhongjun Jin, H. V. JagadishSIGMOD 2021 · 被引用 30 次
- ARDA: Automatic Relational Data Augmentation for Machine LearningNadiia Chepurko, Ryan Marcus, Emanuel Zgraggen, Raul Castro Fernandez 等VLDB 2020 · 被引用 14 次
相关 Paper
- One Size Does Not Fit All: A Bandit-Based Sampler Combination Framework with Theoretical GuaranteesJinglin Peng, Bolin Ding, Jiannan Wang, Kai Zeng 等SIGMOD 2022 · 被引用 7 次
- Approximating a Distribution Using Weight QueriesNadav Barak, Sivan SabatoICML 2021
- CDA: Cost-Sensitive Data Acquisition for Incomplete DatasetsKaiyu Li, Xiaohui Yu, Jian PeiICDE 2025
- Intersectional Fairness in Reinforcement Learning with Large State and Constraint SpacesEric Eaton, Marcel Hussing, Michael Kearns, Aaron Roth 等ICML 2025
- Worst-Case Analysis for Randomly Collected DataJustin Y. Chen, Gregory Valiant, Paul ValiantNeurIPS 2020 · 被引用 4 次
