Assistive Prompt Mediation: Evaluating Language Models Under Accessibility Constraints
Priyaranjan Pattnayak, Ishan Banerjee
Abstract
Large language models (LLMs) are increasingly used as assistive interfaces for users who cannot reliably produce clean text due to accessibility constraints, yet existing evaluations assume iterative input repair and focus on task accuracy or generic noise robustness. We introduce Assistive Prompt Mediation (APM), a theory-grounded evaluation paradigm that reframes assistance as a constrained mediation problem: recovering latent user intent from accessibility-impaired input without clarification, while minimizing cognitive burden and hallucination risk. APM decomposes assistive quality along these axes and is instantiated across 8 languages, 4 accessibility-driven noise classes, and 10 frontier LLMs, with impairment severity yielding accessibility sensitivity curves. Results show that apparent robustness often masks trade-offs—high intent preservation frequently coincides with increased burden or hallucinated mediation, hallucination rates vary by more than across noise types, and assistive decisions exhibit bounded entropy ( normalized), indicating systematic rather than unstable behavior. These findings demonstrate that standard robustness metrics substantially overestimate assistive reliability and motivate evaluating LLMs as constrained mediators under accessibility-driven input degradation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 03f22542-28f8-44fd-8007-702b5c8307d2Builds on5
- Automatic Prompt Optimization with "Gradient Descent" and Beam SearchReid Pryzant, Dan Iter, Jerry Li, Yin Tat Lee et al.EMNLP 2023 · 137 citations
- On Faithfulness and Factuality in Abstractive SummarizationJoshua Maynez, Shashi Narayan, Bernd Bohnet, Ryan T. McDonaldACL 2020 · 54 citations
- HomeSound: An Iterative Field Deployment of an In-Home Sound Awareness System for Deaf or Hard of Hearing UsersDhruv Jain, Kelly Mack, Akli Amrous, Matt Wright et al.CHI 2020 · 52 citations
- AccessEval: Benchmarking Disability Bias in Large Language ModelsSrikant Panda, Amit Agarwal, Hitesh Laxmichand PatelEMNLP 2025 · 2 citations
- Flip-Flop Consistency: Unsupervised Training for Robustness to Prompt Perturbations in LLMsParsa Hejabi, Elnaz Rahmati, Alireza Salkhordeh Ziabari, Morteza DehghaniACL 2026 · 1 citation
Related papers
- Estimating LLM Consistency: A User Baseline vs Surrogate MetricsXiaoyuan Wu, Weiran Lin, Omer Akgul, Lujo BauerEMNLP 2025
- PRISM: Probing Reasoning, Instruction, and Source Memory in LLM HallucinationsYuhe Wu, Guangyu Wang, Yuran Chen, Jiatong Zhang et al.ACL 2026
- Mapping from Meaning: Addressing the Miscalibration of Prompt-Sensitive Language ModelsKyle Cox, Jiawei Xu, Yikun Han, Rong Xu et al.AAAI 2025 · 6 citations
- ResiliBench: Evaluating Agentic Workflow Adaptation in Stochastic EnvironmentsRuicheng Ao, Zeping Min, Tingyu Zhu, Wotao Yin et al.ICLR 2026
- HILL: A Hallucination Identifier for Large Language ModelsFlorian Leiser, Sven Eckhardt, Valentin Leuthe, Merlin Knaeble et al.CHI 2024 · 67 citations
