Assistive Prompt Mediation: Evaluating Language Models Under Accessibility Constraints
Priyaranjan Pattnayak, Ishan Banerjee
摘要
Large language models (LLMs) are increasingly used as assistive interfaces for users who cannot reliably produce clean text due to accessibility constraints, yet existing evaluations assume iterative input repair and focus on task accuracy or generic noise robustness. We introduce Assistive Prompt Mediation (APM), a theory-grounded evaluation paradigm that reframes assistance as a constrained mediation problem: recovering latent user intent from accessibility-impaired input without clarification, while minimizing cognitive burden and hallucination risk. APM decomposes assistive quality along these axes and is instantiated across 8 languages, 4 accessibility-driven noise classes, and 10 frontier LLMs, with impairment severity yielding accessibility sensitivity curves. Results show that apparent robustness often masks trade-offs—high intent preservation frequently coincides with increased burden or hallucinated mediation, hallucination rates vary by more than across noise types, and assistive decisions exhibit bounded entropy ( normalized), indicating systematic rather than unstable behavior. These findings demonstrate that standard robustness metrics substantially overestimate assistive reliability and motivate evaluating LLMs as constrained mediators under accessibility-driven input degradation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Automatic Prompt Optimization with "Gradient Descent" and Beam SearchReid Pryzant, Dan Iter, Jerry Li, Yin Tat Lee 等EMNLP 2023 · 被引用 137 次
- On Faithfulness and Factuality in Abstractive SummarizationJoshua Maynez, Shashi Narayan, Bernd Bohnet, Ryan T. McDonaldACL 2020 · 被引用 54 次
- HomeSound: An Iterative Field Deployment of an In-Home Sound Awareness System for Deaf or Hard of Hearing UsersDhruv Jain, Kelly Mack, Akli Amrous, Matt Wright 等CHI 2020 · 被引用 52 次
- AccessEval: Benchmarking Disability Bias in Large Language ModelsSrikant Panda, Amit Agarwal, Hitesh Laxmichand PatelEMNLP 2025 · 被引用 2 次
- Flip-Flop Consistency: Unsupervised Training for Robustness to Prompt Perturbations in LLMsParsa Hejabi, Elnaz Rahmati, Alireza Salkhordeh Ziabari, Morteza DehghaniACL 2026 · 被引用 1 次
相关 Paper
- Estimating LLM Consistency: A User Baseline vs Surrogate MetricsXiaoyuan Wu, Weiran Lin, Omer Akgul, Lujo BauerEMNLP 2025
- PRISM: Probing Reasoning, Instruction, and Source Memory in LLM HallucinationsYuhe Wu, Guangyu Wang, Yuran Chen, Jiatong Zhang 等ACL 2026
- Mapping from Meaning: Addressing the Miscalibration of Prompt-Sensitive Language ModelsKyle Cox, Jiawei Xu, Yikun Han, Rong Xu 等AAAI 2025 · 被引用 6 次
- ResiliBench: Evaluating Agentic Workflow Adaptation in Stochastic EnvironmentsRuicheng Ao, Zeping Min, Tingyu Zhu, Wotao Yin 等ICLR 2026
- HILL: A Hallucination Identifier for Large Language ModelsFlorian Leiser, Sven Eckhardt, Valentin Leuthe, Merlin Knaeble 等CHI 2024 · 被引用 67 次
