The Lawyer That Never Thinks: Consistency and Fairness as Keys to Reliable AI
Dana R. Alsagheer, Abdulrahman Kamal, Mohammad Kamal, Cosmo Yang Wu, Weidong Shi
摘要
Large Language Models (LLMs) are increasingly used in high-stakes domains like law and research, yet their inconsistencies and response instability raise concerns about trustworthiness. This study evaluates six leading LLMs-GPT-3.5, GPT-4, Claude, Gemini, Mistral, and LLaMA 2-on rationality, stability, and ethical fairness through reasoning tests, legal challenges, and bias-sensitive scenarios. Results reveal significant inconsistencies, highlighting trade-offs between model scale, architecture, and logical coherence. These findings underscore the risks of deploying LLMs in legal and policy settings, emphasizing the need for AI systems that prioritize transparency, fairness, and ethical robustness.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 被引用 3,228 次
- Aligning AI With Shared Human ValuesDan Hendrycks, Collin Burns, Steven Basart, Andrew Critch 等ICLR 2021 · 被引用 878 次
- Surface Form Competition: Why the Highest Probability Answer Isn't Always RightAri Holtzman, Peter West, Vered Shwartz, Yejin Choi 等EMNLP 2021 · 被引用 10 次
相关 Paper
- LLMS ON TRIAL: Evaluating Judicial Fairness For Large Language ModelsYiran Hu, Zongyue Xue, Haitao Li, Siyuan Zheng 等ICLR 2026 · 被引用 4 次
- On the Reliability of Psychological Scales on Large Language ModelsJen-tse Huang, Wenxiang Jiao, Man Ho Lam, Eric John Li 等EMNLP 2024 · 被引用 6 次
- Nuance Matters: Probing Epistemic Consistency in Causal ReasoningShaobo Cui, Junyou Li, Luca Mouchel, Yiyang Feng 等AAAI 2025 · 被引用 2 次
- Measuring the Unmeasurable: Unveiling Latent Cognitive Capabilities of LLMCui Danxin, Sihang Jiang, Keyi Wang, Zhiyi Duan 等AAAI 2026
- A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and RecommendationsMd. Tahmid Rahman Laskar, Sawsan Alqahtani, M. Saiful Bari, Mizanur Rahman 等EMNLP 2024 · 被引用 47 次
