Bias Neutralization in Non-Parallel Texts: A Cyclic Approach with Auxiliary Guidance
Karthic Madanagopal, James Caverlee
摘要
Objectivity is a goal for Wikipedia and many news sites, as well as a guiding principle of many large language models. Indeed, several methods have recently been developed for automatic subjective bias neutralization. These methods, however, typically rely on parallel text for training (i.e. a biased sentence coupled with a non-biased sentence), demonstrate poor transfer to new domains, and can lose important bias-independent context. Toward expanding the reach of bias neutralization, we propose in this paper a new approach called FairBalance. Three of its unique features are: i) a cycle consistent adversarial network enables bias neutralization without the need for parallel text; ii) the model design preserves bias-independent content; and iii) through auxiliary guidance, the model highlights sequences of bias-inducing words, yielding strong results in terms of bias neutralization quality. In our evaluations involving seven models comprising of adversarial and non-adversarial models, the FairBalance method showed a notable improvement in bias neutralization based on subjective human judgment when compared to other techniques.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Hate Speech Detection with Generalizable Target-aware FairnessTong Chen, Danny Wang, Xurong Liang, Marten Risius 等KDD 2024 · 被引用 7 次
- FairFlow: Mitigating Dataset Biases through Undecided Learning for Natural Language UnderstandingJiali Cheng, Hadi AmiriEMNLP 2024 · 被引用 2 次
它引用的顶会 Paper7
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- A Probabilistic Formulation of Unsupervised Text Style TransferJunxian He, Xinyi Wang, Graham Neubig, Taylor Berg-KirkpatrickICLR 2020 · 被引用 136 次
- Hooks in the Headline: Learning to Generate Headlines with Controlled StylesDi Jin, Zhijing Jin, Joey Tianyi Zhou, Lisa Orii 等ACL 2020 · 被引用 56 次
- Revision in Continuous Space: Unsupervised Text Style Transfer without Adversarial LearningDayiheng Liu, Jie Fu, Yidan Zhang, Chris Pal 等AAAI 2020 · 被引用 53 次
- A Transformer-based Framework for Neutralizing and Reversing the Political Polarity of News ArticlesRuibo Liu, Chenyan Jia, Soroush VosoughiCSCW 2021 · 被引用 25 次
相关 Paper
- KLAAD: Refining Attention Mechanisms to Reduce Societal Bias in Generative Language ModelsSeorin Kim, Dongyoung Lee, Jaejin LeeEMNLP 2025
- Mitigating Biases in Language Models via Bias UnlearningDianqing Liu, Yi Liu, Guoqing Jin, Zhendong MaoEMNLP 2025 · 被引用 4 次
- FairFil: Contrastive Neural Debiasing Method for Pretrained Text EncodersPengyu Cheng, Weituo Hao, Siyang Yuan, Shijing Si 等ICLR 2021 · 被引用 50 次
- Auto-Search and Refinement: An Automated Framework for Gender Bias Mitigation in Large Language ModelsYue Xu, Chengyan Fu, Li Xiong, Sibei Yang 等NeurIPS 2025 · 被引用 3 次
- InvDiff: Invariant Guidance for Bias Mitigation in Diffusion ModelsMin Hou, Yueying Wu, Chang Xu, Yu-Hao Huang 等KDD 2025 · 被引用 2 次
