Bias Neutralization in Non-Parallel Texts: A Cyclic Approach with Auxiliary Guidance
Karthic Madanagopal, James Caverlee
Abstract
Objectivity is a goal for Wikipedia and many news sites, as well as a guiding principle of many large language models. Indeed, several methods have recently been developed for automatic subjective bias neutralization. These methods, however, typically rely on parallel text for training (i.e. a biased sentence coupled with a non-biased sentence), demonstrate poor transfer to new domains, and can lose important bias-independent context. Toward expanding the reach of bias neutralization, we propose in this paper a new approach called FairBalance. Three of its unique features are: i) a cycle consistent adversarial network enables bias neutralization without the need for parallel text; ii) the model design preserves bias-independent content; and iii) through auxiliary guidance, the model highlights sequences of bias-inducing words, yielding strong results in terms of bias neutralization quality. In our evaluations involving seven models comprising of adversarial and non-adversarial models, the FairBalance method showed a notable improvement in bias neutralization based on subjective human judgment when compared to other techniques.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Hate Speech Detection with Generalizable Target-aware FairnessTong Chen, Danny Wang, Xurong Liang, Marten Risius et al.KDD 2024 · 7 citations
- FairFlow: Mitigating Dataset Biases through Undecided Learning for Natural Language UnderstandingJiali Cheng, Hadi AmiriEMNLP 2024 · 2 citations
Builds on7
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- A Probabilistic Formulation of Unsupervised Text Style TransferJunxian He, Xinyi Wang, Graham Neubig, Taylor Berg-KirkpatrickICLR 2020 · 136 citations
- Hooks in the Headline: Learning to Generate Headlines with Controlled StylesDi Jin, Zhijing Jin, Joey Tianyi Zhou, Lisa Orii et al.ACL 2020 · 56 citations
- Revision in Continuous Space: Unsupervised Text Style Transfer without Adversarial LearningDayiheng Liu, Jie Fu, Yidan Zhang, Chris Pal et al.AAAI 2020 · 53 citations
- A Transformer-based Framework for Neutralizing and Reversing the Political Polarity of News ArticlesRuibo Liu, Chenyan Jia, Soroush VosoughiCSCW 2021 · 25 citations
Related papers
- KLAAD: Refining Attention Mechanisms to Reduce Societal Bias in Generative Language ModelsSeorin Kim, Dongyoung Lee, Jaejin LeeEMNLP 2025
- Mitigating Biases in Language Models via Bias UnlearningDianqing Liu, Yi Liu, Guoqing Jin, Zhendong MaoEMNLP 2025 · 4 citations
- FairFil: Contrastive Neural Debiasing Method for Pretrained Text EncodersPengyu Cheng, Weituo Hao, Siyang Yuan, Shijing Si et al.ICLR 2021 · 50 citations
- Auto-Search and Refinement: An Automated Framework for Gender Bias Mitigation in Large Language ModelsYue Xu, Chengyan Fu, Li Xiong, Sibei Yang et al.NeurIPS 2025 · 3 citations
- InvDiff: Invariant Guidance for Bias Mitigation in Diffusion ModelsMin Hou, Yueying Wu, Chang Xu, Yu-Hao Huang et al.KDD 2025 · 2 citations
