FairLex: A Multilingual Benchmark for Evaluating Fairness in Legal Text Processing
Ilias Chalkidis, Tommaso Pasini, Sheng Zhang, Letizia Tomada, Sebastian Felix Schwemer, Anders Søgaard
Abstract
We present a benchmark suite of four datasets for evaluating the fairness of pre-trained language models and the techniques used to fine-tune them for downstream tasks. Our benchmarks cover four jurisdictions (European Council, USA, Switzerland, and China), five languages (English, German, French, Italian and Chinese) and fairness across five attributes (gender, age, region, language, and legal area). In our experiments, we evaluate pretrained language models using several grouprobust fine-tuning techniques and show that performance group disparities are vibrant in many cases, while none of these techniques guarantee fairness, nor consistently mitigate group disparities. Furthermore, we provide a quantitative and qualitative analysis of our results, highlighting open challenges in the development of robustness methods in legal NLP.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b40c6b3a-96f3-4721-bf7b-e3be9ea3002aCited by top-tier papers12
- Aging with GRACE: Lifelong Model Editing with Discrete Key-Value AdaptorsTom Hartvigsen, Swami Sankaranarayanan, Hamid Palangi, Yoon Kim et al.NeurIPS 2023 · 349 citations
- MELO: Enhancing Model Editing with Neuron-Indexed Dynamic LoRALang Yu, Qin Chen, Jie Zhou, Liang HeAAAI 2024 · 96 citations
- On Learning Fairness and Accuracy on Multiple SubgroupsChangjian Shui, Gezheng Xu, Qi Chen, Jiaqi Li et al.NeurIPS 2022 · 58 citations
- LeXFiles and LegalLAMA: Facilitating English Multinational Legal Language Model DevelopmentIlias Chalkidis, Nicolas Garneau, Catalina Goanta, Daniel Martin Katz et al.ACL 2023 · 29 citations
- Deconfounding Legal Judgment Prediction for European Court of Human Rights Cases Towards Better Alignment with ExpertsTokala Yaswanth Sri Sai Santosh, Shanshan Xu, Oana Ichim, Matthias GrabmairEMNLP 2022 · 13 citations
Builds on7
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- Distributionally Robust Neural NetworksShiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, Percy LiangICLR 2020 · 1,578 citations
- Out-of-Distribution Generalization via Risk Extrapolation (REx)David Krueger, Ethan Caballero, Jörn-Henrik Jacobsen, Amy Zhang et al.ICML 2021 · 1,163 citations
- Retiring Adult: New Datasets for Fair Machine LearningFrances Ding, Moritz Hardt, John Miller, Ludwig SchmidtNeurIPS 2021 · 671 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
Related papers
- Fairness Beyond Performance: Revealing Reliability Disparities Across Groups in Legal NLPT. Y. S. S. Santosh, Irtiza ChowdhuryACL 2025
- EuroGEST: Investigating gender stereotypes in multilingual language modelsJacqueline Rowe, Mateusz Klimaszewski, Liane Guillou, Shannon Vallor et al.EMNLP 2025
- Auto-Debias: Debiasing Masked Language Models with Automated Biased PromptsYue Guo, Yi Yang, Ahmed AbbasiACL 2022
- Equi-Tuning: Group Equivariant Fine-Tuning of Pretrained ModelsSourya Basu, Prasanna Sattigeri, Karthikeyan Natesan Ramamurthy, Vijil Chenthamarakshan et al.AAAI 2023 · 25 citations
- Perturbation Augmentation for Fairer NLPRebecca Qian, Candace Ross, Jude Fernandes, Eric Michael Smith et al.EMNLP 2022 · 54 citations
