Causal-Guided Active Learning for Debiasing Large Language Models
Zhouhao Sun, Li Du, Xiao Ding, Yixuan Ma, Yang Zhao, Kaitao Qiu, Ting Liu, Bing Qin
Abstract
Although achieving promising performance, recent analyses show that current generative large language models (LLMs) may still capture dataset biases and utilize them for generation, leading to poor generalizability and harmfulness of LLMs. However, due to the diversity of dataset biases and the over-optimization problem, previous prior-knowledge-based debiasing methods and fine-tuning-based debiasing methods may not be suitable for current LLMs. To address this issue, we explore combining active learning with the causal mechanisms and propose a casual-guided active learning (CAL) framework, which utilizes LLMs itself to automatically and autonomously identify informative biased samples and induce the bias patterns. Then a cost-effective and efficient incontext learning based method is employed to prevent LLMs from utilizing dataset biases during generation. Experimental results show that CAL can effectively recognize typical biased instances and induce various bias patterns for debiasing LLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eb1184bb-92ec-438f-9142-a2727d126de3Cited by top-tier papers9
- From Selection to Generation: A Survey of LLM-based Active LearningYu Xia, Subhojyoti Mukherjee, Zhouhang Xie, Junda Wu et al.ACL 2025 · 18 citations
- Lying with Truths: Open-Channel Multi-Agent Collusion for Belief Manipulation via Generative MontageJinwei Hu, Xinmiao Huang, Youcheng Sun, Yi Dong et al.ACL 2026 · 11 citations
- PIPER: Benchmarking and Prompting Event Reasoning Boundary of LLMs via Debiasing-Distillation Enhanced TuningZhicong Lu, Changyuan Tian, PeiguangLi PeiguangLi, Li Jin et al.ACL 2025 · 4 citations
- AgriEval: A Comprehensive Chinese Agricultural Benchmark for Large Language ModelsLian Yan, Haotian Wang, Chen Tang, Haifeng Liu et al.AAAI 2026 · 3 citations
- Data Selection Matters: Towards Robust Instruction Tuning of Large Multimodal ModelsXu Yang, Chen Liu, Ying WeiNeurIPS 2025 · 2 citations
Builds on12
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal et al.ACL 2020 · 602 citations
- ExT5: Towards Extreme Multi-Task Scaling for Transfer LearningVamsi Aribandi, Yi Tay, Tal Schuster, Jinfeng Rao et al.ICLR 2022 · 237 citations
- An LLM can Fool Itself: A Prompt-Based Adversarial AttackXilie Xu, Keyi Kong, Ning Liu, Lizhen Cui et al.ICLR 2024 · 146 citations
- Learning from others' mistakes: Avoiding dataset biases without modeling themVictor Sanh, Thomas Wolf, Yonatan Belinkov, Alexander M. RushICLR 2021 · 123 citations
- Prompting GPT-3 To Be ReliableChenglei Si, Zhe Gan, Zhengyuan Yang, Shuohang Wang et al.ICLR 2023 · 68 citations
Related papers
- Prompting Fairness: Integrating Causality to Debias Large Language ModelsJingling Li, Zeyu Tang, Xiaoyu Liu, Peter Spirtes et al.ICLR 2025
- A Causal Explainable Guardrails for Large Language ModelsZhixuan Chu, Yan Wang, Longfei Li, Zhibo Wang et al.CCS 2024 · 5 citations
- Causal Prompting: Debiasing Large Language Model Prompting Based on Front-Door AdjustmentCongzhi Zhang, Linhai Zhang, Jialong Wu, Yulan He et al.AAAI 2025 · 42 citations
- Knowing Bias, Doing Better: Mitigating Social Bias in LLMs via Know-Bias Neuron EnhancementJinhao Pan, Chahat Raj, Anjishnu Mukherjee, Sina Mansouri et al.ICML 2026
- CCL: Causal-aware In-context Learning for Out-of-Distribution GeneralizationHoyoon Byun, Gyeongdeok Seo, Joonseong Kang, Taero Kim et al.NeurIPS 2025 · 1 citation
