AccessEval: Benchmarking Disability Bias in Large Language Models
Srikant Panda, Amit Agarwal, Hitesh Laxmichand Patel
Abstract
Large Language Models (LLMs) are increasingly deployed across diverse domains but often exhibit disparities in how they handle real-life queries. To systematically investigate these effects within various disability contexts, we introduce AccessEval (Accessibility Evaluation), a benchmark evaluating 21 closed- and open-source LLMs across 6 real-world domains and 9 disability types using paired Neutral and Disability-Aware Queries. We evaluated model outputs with metrics for sentiment, social perception, and factual accuracy. Our analysis reveals that responses to disability-aware queries tend to have a more negative tone, increased stereotyping, and higher factual error compared to neutral queries. These effects show notable variation by domain and disability type, with disabilities affecting hearing, speech, and mobility disproportionately impacted. These disparities reflect persistent forms of ableism embedded in model behavior. By examining model performance in real-world decision-making contexts, we better illuminate how such biases can translate into tangible harms for disabled users. This framing helps bridges the gap between technical evaluation and user impact, reinforcing importance of bias mitigation in day-to-day applications. Our dataset is publicly available at: https://huggingface.co/datasets/Srikant86/AccessEval
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 91efd4d4-ae16-4dfa-b601-302081314d80Cited by top-tier papers2
- Assistive Prompt Mediation: Evaluating Language Models Under Accessibility ConstraintsPriyaranjan Pattnayak, Ishan BanerjeeICML 2026
- Emergence of Biased Consensus in Multi-Agent LLM DebatesMaya OkawaICML 2026
Builds on4
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Climbing towards NLU: On Meaning, Form, and Understanding in the Age of DataEmily M. Bender, Alexander KollerACL 2020 · 914 citations
- LayoutLM: Pre-training of Text and Layout for Document Image UnderstandingYiheng Xu, Minghao Li, Lei Cui, Shaohan Huang et al.KDD 2020 · 575 citations
- StereoSet: Measuring stereotypical bias in pretrained language modelsMoin Nadeem, Anna Bethke, Siva ReddyACL 2021
Related papers
- BiasFreeBench: a Benchmark for Mitigating Bias in Large Language Model ResponsesXin Xu, Xunzhi He, Churan Zhi, Ruizhe Chen et al.ICLR 2026 · 4 citations
- Social Bias Probing: Fairness Benchmarking for Language ModelsMarta Marchiori Manerba, Karolina Stanczak, Riccardo Guidotti, Isabelle AugensteinEMNLP 2024 · 4 citations
- Who is better at math, Jenny or Jingzhen? Uncovering Stereotypes in Large Language ModelsZara Siddique, Liam D. Turner, Luis Espinosa AnkeEMNLP 2024 · 2 citations
- Adaptive Generation of Bias-Eliciting Questions for LLMsRobin Staab, Jasper Dekoninck, Maximilian Baader, Martin VechevICML 2026
- Identifying, Explaining, and Correcting Ableist Language with AIKynnedy Simone Smith, Lydia B. Chilton, Danielle BraggCHI 2026 · 1 citation
