Same Same, But Different: Conditional Multi-Task Learning for Demographic-Specific Toxicity Detection
Soumyajit Gupta, Sooyong Lee, Maria De-Arteaga, Matthew Lease
摘要
Algorithmic bias often arises as a result of differential subgroup validity, in which predictive relationships vary across groups. For example, in toxic language detection, comments targeting different demographic groups can vary markedly across groups. In such settings, trained models can be dominated by the relationships that best fit the majority group, leading to disparate performance. We propose framing toxicity detection as multi-task learning (MTL), allowing a model to specialize on the relationships that are relevant to each demographic group while also leveraging shared properties across groups. With toxicity detection, each task corresponds to identifying toxicity against a particular demographic group. However, traditional MTL requires labels for all tasks to be present for every data point. To address this, we propose Conditional MTL (CondMTL), wherein only training examples relevant to the given demographic group are considered by the loss function. This lets us learn group specific representations in each branch which are not cross contaminated by irrelevant labels. Results on synthetic and real data show that using CondMTL improves predictive recall over various baselines in general and for the minority demographic group in particular, while having similar overall accuracy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Human-centered NLP Fact-checking: Co-Designing with Fact-checkers using Matchmaking for AIHoujiang Liu, Anubrata Das, Alexander Boltz, Didi Zhou 等CSCW 2024 · 被引用 23 次
- Hate Speech Detection with Generalizable Target-aware FairnessTong Chen, Danny Wang, Xurong Liang, Marten Risius 等KDD 2024 · 被引用 7 次
- Voices in a Crowd: Searching for clusters of unique perspectivesNikolas Vitsakis, Amit Parekh, Ioannis KonstasEMNLP 2024
它引用的顶会 Paper6
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie 等ICML 2021 · 被引用 1,773 次
- FairBatch: Batch Selection for Model FairnessYuji Roh, Kangwook Lee, Steven Euijong Whang, Changho SuhICLR 2021 · 被引用 156 次
- Jury Learning: Integrating Dissenting Voices into Machine Learning ModelsMitchell L. Gordon, Michelle S. Lam, Joon Sung Park, Kayur Patel 等CHI 2022 · 被引用 134 次
- Many Task Learning With Task RoutingGjorgji Strezoski, Nanne van Noord, Marcel WorringICCV 2019 · 被引用 112 次
- Fragile Masculinity: Men, Gender, and Online HarassmentJennifer D. Rubin, Lindsay Blackwell, Terri D. ConleyCHI 2020 · 被引用 40 次
相关 Paper
- Bias Mitigation for Toxicity Detection via Sequential DecisionsLu Cheng, Ahmadreza Mosallanezhad, Yasin N. Silva, Deborah L. Hall 等SIGIR 2022 · 被引用 8 次
- Learning Subjective Label Distributions via Sociocultural DescriptorsMohammed Fayiz Parappan, Ricardo HenaoEMNLP 2025 · 被引用 5 次
- Conditional Learning of Fair RepresentationsHan Zhao, Amanda Coston, Tameem Adel, Geoffrey J. GordonICLR 2020 · 被引用 127 次
- Exploiting MMD and Sinkhorn Divergences for Fair and Transferable Representation LearningLuca Oneto, Michele Donini, Giulia Luise, Carlo Ciliberto 等NeurIPS 2020 · 被引用 56 次
- Multi-Task Representation Alignment on Language Understanding: A Mutual Information PerspectiveDou Hu, Lingwei Wei, Hongjiang Xiao, Songlin Hu 等ACL 2026
