CHBias: Bias Evaluation and Mitigation of Chinese Conversational Language Models
Jiaxu Zhao, Meng Fang, Zijing Shi, Yitong Li, Ling Chen, Mykola Pechenizkiy
Abstract
redWarning: This paper contains content that may be offensive or upsetting.Pretrained conversational agents have been exposed to safety issues, exhibiting a range of stereotypical human biases such as gender bias. However, there are still limited bias categories in current research, and most of them only focus on English. In this paper, we introduce a new Chinese dataset, CHBias, for bias evaluation and mitigation of Chinese conversational language models.Apart from those previous well-explored bias categories, CHBias includes under-explored bias categories, such as ageism and appearance biases, which received less attention. We evaluate two popular pretrained Chinese conversational models, CDial-GPT and EVA2.0, using CHBias. Furthermore, to mitigate different biases, we apply several debiasing methods to the Chinese pretrained models. Experimental results show that these Chinese pretrained models are potentially risky for generating texts that contain social biases, and debiasing methods using the proposed dataset can make response generation less biased while preserving the models' conversational capabilities.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Understanding Large Language Model Vulnerabilities to Social Bias AttacksJiaxu Zhao, Meng Fang, Fanghua Ye, Ke Xu et al.ACL 2025
- Incorporating Diverse Perspectives in Cultural Alignment: Survey of Evaluation Benchmarks Through A Three-Dimensional FrameworkMeng-Chen Wu, Si-Chi Chin, Tess Wood, Ayush Goyal et al.EMNLP 2025
- GeWu: A Culturally-Grounded Chinese Benchmark for Multi-Stage Social Bias Evaluation in Large Language ModelsYi Lin, Ziyi Zhou, Jiashi Gao, Xinwei Guo et al.AAAI 2026
- BTC-SAM: Leveraging LLMs for Generation of Bias Test Cases for Sentiment Analysis ModelsZsolt T. Kardkovács, Lynda Djennane, Anna Field, Boualem Benatallah et al.EMNLP 2025
Builds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- PLATO: Pre-trained Dialogue Generation Model with Discrete Latent VariableSiqi Bao, Huang He, Fan Wang, Hua Wu et al.ACL 2020 · 229 citations
- Towards Conversational Recommendation over Multi-Type DialogsZeming Liu, Haifeng Wang, Zheng-Yu Niu, Hua Wu et al.ACL 2020 · 157 citations
- KdConv: A Chinese Multi-domain Dialogue Dataset Towards Multi-turn Knowledge-driven ConversationHao Zhou, Chujie Zheng, Kaili Huang, Minlie Huang et al.ACL 2020 · 106 citations
- Mitigating Gender Bias for Neural Dialogue Generation with Adversarial LearningHaochen Liu, Wentao Wang, Yiqi Wang, Hui Liu et al.EMNLP 2020 · 55 citations
Related papers
- RedditBias: A Real-World Resource for Bias Evaluation and Debiasing of Conversational Language ModelsSoumya Barikeri, Anne Lauscher, Ivan Vulic, Goran GlavasACL 2021
- Auto-Debias: Debiasing Masked Language Models with Automated Biased PromptsYue Guo, Yi Yang, Ahmed AbbasiACL 2022
- Social Debiasing for Fair Multi-Modal LLMsHarry Cheng, Yangyang Guo, Qing Guo, Ming-Hsuan Yang et al.ICCV 2025 · 1 citation
- Causal-Debias: Unifying Debiasing in Pretrained Language Models and Fine-tuning via Causal Invariant LearningFan Zhou, Yuzhou Mao, Liu Yu, Yi Yang et al.ACL 2023 · 21 citations
- BERTScore is Unfair: On Social Bias in Language Model-Based Metrics for Text GenerationTianxiang Sun, Junliang He, Xipeng Qiu, Xuanjing HuangEMNLP 2022 · 22 citations
