Profiling the Irrational Agent: Cognitive Modeling of LLM Behaviors in Sequential Jailbreaks
Xikang Yang, Biyu Zhou, Xuehai Tang, Jizhong Han, Songlin Hu
Abstract
Large language models (LLMs) are increasingly deployed in high-stakes settings, yet they remain vulnerable to sequential jailbreaks that exploit multi-turn interaction to circumvent safety mechanisms. Current safety evaluations are largely outcome-based, offering little insight into the latent decision processes that lead to unsafe compliance. We propose an interpretable cognitive modeling framework that couples a controlled elicitation paradigm, the Contextual Iowa Gambling Task (C-IGT), with a Generalized Rescorla--Wagner (GRW) architecture to decompose behavior into measurable mechanisms. Across a diverse set of mainstream LLMs, we find that sequential vulnerability is not explained by scale alone but emerges from interactions among cognitive factors, including optimism-biased learning, perceptual reward amplification, and choice inertia. Moreover, counterfactual feedback and psychologically framed rewards (e.g., regret, authority, threat) substantially accelerate the transition from refusal to compliance. These results yield principled cognitive profiles of LLM "irrationality" and provide insights for interdisciplinary research on LLM agents at the intersection of machine learning and human behavioral science.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 174be4c9-be2b-4182-be20-c1816274ba9fBuilds on10
- Tree of Attacks: Jailbreaking Black-Box LLMs AutomaticallyAnay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson et al.NeurIPS 2024 · 835 citations
- How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMsYi Zeng, Hongpeng Lin, Jingwen Zhang, Diyi Yang et al.ACL 2024 · 64 citations
- CogBench: a large language model walks into a psychology labJulian Coda-Forno, Marcel Binz, Jane X. Wang, Eric SchulzICML 2024 · 60 citations
- ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMsFengqing Jiang, Zhangchen Xu, Luyao Niu, Zhen Xiang et al.ACL 2024 · 36 citations
- In-Context Learning Agents Are Asymmetric Belief UpdatersJohannes A. Schubert, Akshay K. Jagadish, Marcel Binz, Eric SchulzICML 2024 · 16 citations
Related papers
- Using Reinforcement Learning to Train Large Language Models to Explain Human DecisionsJian-Qiao Zhu, Hanbo Xie, Dilip Arumugam, Robert C. Wilson et al.ICLR 2026 · 10 citations
- Exploiting Synergistic Cognitive Biases to Bypass Safety in LLMsXikang Yang, Biyu Zhou, Xuehai Tang, Jizhong Han et al.AAAI 2026
- Adjacent Words, Divergent Intents: Jailbreaking Large Language Models via Task ConcurrencyYukun Jiang, Mingjie Li, Michael Backes, Yang ZhangNeurIPS 2025 · 17 citations
- Accommodation and Epistemic Vigilance: A Pragmatic Account of Why LLMs Fail to Challenge Harmful BeliefsMyra Cheng, Robert D. Hawkins, Dan JurafskyACL 2026 · 6 citations
- Cannot See the Forest for the Trees: Invoking Heuristics and Biases to Elicit Irrational Choices of LLMsHaoming Yang, Ke Ma, Xiaojun Jia, Yingfei Sun et al.ICML 2025
