Towards Fair Truth Discovery from Biased Crowdsourced Answers
Yanying Li, Haipei Sun, Wendy Hui Wang
Abstract
Crowdsourcing systems have gained considerable interest and adoption in recent years. One important research problem for crowdsourcing systems is truth discovery, which aims to aggregate noisy answers contributed by the workers to obtain the correct answer (truth) of each task. However, since the collected answers are highly prone to the workers' biases, aggregating these biased answers without proper treatment will unavoidably lead to discriminatory truth discovery results for particular race, gender and political groups. To address this challenge, in this paper, first, we define a new fairness notion named θ-disparity for truth discovery. Intuitively, θ-disparity bounds the difference in the probabilities that the truth of both protected and unprotected groups being predicted to be positive. Second, we design three fairness enhancing methods, namely Pre-TD, FairTD, and Post-TD, for truth discovery. Pre-TD is a pre-processing method that removes the bias in workers' answers before truth discovery. FairTD is an in-processing method that incorporates fairness into the truth discovery process. And Post-TD is a post-processing method that applies additional treatment on the discovered truth to make it satisfy θ-disparity. We perform an extensive set of experiments on both synthetic and real-world crowdsourcing datasets. Our results demonstrate that among the three fairness enhancing methods, FairTD produces the best accuracy with θ-disparity. In some settings, the accuracy of FairTD is even better than truth discovery without fairness, as it removes some low-quality answers as side effects.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers6
- Crowdsourcing Subjective Annotations Using Pairwise Comparisons Reduces Bias and Error Compared to the Majority-vote MethodHasti Narimanzadeh, Arash Badie Modiri, Iuliia G. Smirnova, Ted Hsuan Yun ChenCSCW 2023 · 20 citations
- Chameleon: Foundation Models for Fairness-aware Multi-modal Data Augmentation to Enhance Coverage of MinoritiesMahdi Erfanian, H. V. Jagadish, Abolfazl AsudehVLDB 2024 · 10 citations
- Product Question Answering in E-Commerce: A SurveyYang Deng, Wenxuan Zhang, Qian Yu, Wai LamACL 2023 · 9 citations
- Noisy Interactive Graph SearchQianhao Cong, Jing Tang, Kai Han, Yuming Huang et al.KDD 2022 · 4 citations
- Optimal Fair Aggregation of Crowdsourced Noisy Labels using Demographic Parity ConstraintsGabriel Singer, Samuel Gruffaz, Olivier VO VAN, Nicolas Vayatis et al.ICML 2026
Related papers
- RCTD: Reputation-Constrained Truth Discovery in Sybil Attack Crowdsourcing EnvironmentXing Jin, Zhihai Gong, Jiuchuan Jiang, Chao Wang et al.KDD 2024 · 2 citations
- Towards Personalized Privacy-Preserving Incentive for Truth Discovery in Crowdsourced Binary-Choice Question AnsweringPeng Sun, Zhibo Wang, Yunhe Feng, Liantao Wu et al.INFOCOM 2020 · 39 citations
- Frustratingly Easy Truth DiscoveryReshef Meir, Ofra Amir, Omer Ben-Porat, Tsviel Ben Shabat et al.AAAI 2023 · 2 citations
- Can The Crowd Identify Misinformation Objectively?: The Effects of Judgment Scale and Assessor's BackgroundKevin Roitero, Michael Soprano, Shaoyang Fan, Damiano Spina et al.SIGIR 2020 · 2 citations
- Fair Sequential Selection Using Supervised Learning ModelsMohammad Mahdi Khalili, Xueru Zhang, Mahed AbroshanNeurIPS 2021 · 25 citations
