The Lawyer That Never Thinks: Consistency and Fairness as Keys to Reliable AI
Dana R. Alsagheer, Abdulrahman Kamal, Mohammad Kamal, Cosmo Yang Wu, Weidong Shi
Abstract
Large Language Models (LLMs) are increasingly used in high-stakes domains like law and research, yet their inconsistencies and response instability raise concerns about trustworthiness. This study evaluates six leading LLMs-GPT-3.5, GPT-4, Claude, Gemini, Mistral, and LLaMA 2-on rationality, stability, and ethical fairness through reasoning tests, legal challenges, and bias-sensitive scenarios. Results reveal significant inconsistencies, highlighting trade-offs between model scale, architecture, and logical coherence. These findings underscore the risks of deploying LLMs in legal and policy settings, emphasizing the need for AI systems that prioritize transparency, fairness, and ethical robustness.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0df02018-4d55-468b-8222-9ee334ea3facBuilds on5
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- Aligning AI With Shared Human ValuesDan Hendrycks, Collin Burns, Steven Basart, Andrew Critch et al.ICLR 2021 · 878 citations
- Surface Form Competition: Why the Highest Probability Answer Isn't Always RightAri Holtzman, Peter West, Vered Shwartz, Yejin Choi et al.EMNLP 2021 · 10 citations
Related papers
- LLMS ON TRIAL: Evaluating Judicial Fairness For Large Language ModelsYiran Hu, Zongyue Xue, Haitao Li, Siyuan Zheng et al.ICLR 2026 · 4 citations
- On the Reliability of Psychological Scales on Large Language ModelsJen-tse Huang, Wenxiang Jiao, Man Ho Lam, Eric John Li et al.EMNLP 2024 · 6 citations
- Nuance Matters: Probing Epistemic Consistency in Causal ReasoningShaobo Cui, Junyou Li, Luca Mouchel, Yiyang Feng et al.AAAI 2025 · 2 citations
- Measuring the Unmeasurable: Unveiling Latent Cognitive Capabilities of LLMCui Danxin, Sihang Jiang, Keyi Wang, Zhiyi Duan et al.AAAI 2026
- A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and RecommendationsMd. Tahmid Rahman Laskar, Sawsan Alqahtani, M. Saiful Bari, Mizanur Rahman et al.EMNLP 2024 · 47 citations
