Language Models Can Predict Their Own Behavior
Dhananjay Ashok, Jonathan May
Abstract
The text produced by language models (LMs) can exhibit specific 'behaviors,' such as a failure to follow alignment training, that we hope to detect and react to during deployment. Identifying these behaviors can often only be done post facto, i.e., after the entire text of the output has been generated. We provide evidence that there are times when we can predict how an LM will behave early in computation, before even a single token is generated. We show that probes trained on the internal representation of input tokens alone can predict a wide range of eventual behaviors over the entire output sequence. Using methods from conformal prediction, we provide provable bounds on the estimation error of our probes, creating precise early warning systems for these behaviors. The conformal probes can identify instances that will trigger alignment failures (jailbreaking) and instruction-following failures, without requiring a single token to be generated. An early warning system built on the probes reduces jailbreaking by 91%. Our probes also show promise in pre-emptively estimating how confident the model will be in its response, a behavior that cannot be detected using the output text alone. Conformal probes can preemptively estimate the final prediction of an LM that uses Chain-of-Thought (CoT) prompting, hence accelerating inference. When applied to an LM that uses CoT to perform text classification, the probes drastically reduce inference costs (65% on average across 27 datasets), with negligible accuracy loss. Encouragingly, probes generalize to unseen datasets and perform better on larger models, suggesting applicability to the largest of models in real-world settings. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8a8080cf-bf7b-4530-9902-988c2555260aCited by top-tier papers4
- Reasoning Theater: Disentangling Model Beliefs from Chain-of-ThoughtSiddharth Boppana, Annabel Ma, Max Loeffler, Raphaël Sarfati et al.ICML 2026 · 30 citations
- Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM SafetySeongmin Lee, Aeree Cho, Grace C. Kim, Shengyun Peng et al.EMNLP 2025 · 1 citation
- Less is More: Geometric Unlearning for LLMs with Minimal Data DisclosureChenchen Tan, Xinghao Li, Shujie Cui, Youyang Qu et al.ICML 2026
- From Input Perception to Predictive Insight: Modeling Model Blind Spots Before They Become ErrorsMaggie Mi, Aline Villavicencio, Nafise Sadat MoosaviEMNLP 2025
Builds on27
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
Related papers
- Catastrophic Jailbreak of Open-source LLMs via Exploiting GenerationYangsibo Huang, Samyak Gupta, Mengzhou Xia, Kai Li et al.ICLR 2024 · 481 citations
- "I've Decided to Leak": Probing Internals Behind Prompt Leakage IntentsJianshuo Dong, Yutong Zhang, Liu Yan, Zhenyu Zhong et al.EMNLP 2025 · 1 citation
- Defending Jailbreak Attacks on Large Language Models via Manifold Trajectory KineticsHangtao Zhang, Yucheng Zhao, Sishun Liu, Ziqi Zhou et al.USENIX Security 2026 · 3 citations
- Quantifying Large Language Model Attacks Through the Lens of Model CognitionXiuming Liu, Chaoxiang He, Xuanran Yu, Jichen Chai et al.USENIX Security 2026
- Mission Impossible: A Statistical Perspective on Jailbreaking LLMsJingtong Su, Julia Kempe, Karen UllrichNeurIPS 2024 · 38 citations
