All Code, No Thought: Language Models Struggle to Reason in Ciphered Language
Shiyuan Guo, Henry Sleight, Fabien Roger
摘要
Detecting harmful AI actions is important as AI agents gain adoption. Chain-of-thought (CoT) monitoring is one method widely used to detect adversarial attacks and AI misalignment. However, attackers and misaligned models might evade CoT monitoring through ciphered reasoning: reasoning hidden in encrypted, translated, or compressed text. To assess this risk, we test whether models can perform ciphered reasoning. For each of 28 different ciphers, we fine-tune and prompt up to 10 models to reason in that cipher. We measure model accuracy on math problems as a proxy for reasoning ability. Across the models we test, we find an asymmetry: model accuracy can drop significantly when reasoning in ciphered text, even though models demonstrate comprehension of ciphered text by being able to translate it accurately to English. Even frontier models struggle with lesser-known ciphers, although they can reason accurately in well-known ciphers like rot13. We show that ciphered reasoning capability correlates with cipher prevalence in pretraining data. We also identify scaling laws showing that ciphered reasoning capability improves slowly with additional fine-tuning data. Our work suggests that evading CoT monitoring using ciphered reasoning may be an ineffective tactic for current models and offers guidance on constraining the development of this capability in future frontier models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen 等ICLR 2024 · 被引用 1,104 次
- Covert Malicious Finetuning: Challenges in Safeguarding LLM AdaptationDanny Halawi, Alexander Wei, Eric Wallace, Tony Tong Wang 等ICML 2024 · 被引用 77 次
- Jailbreaking Large Language Models Against Moderation Guardrails via Cipher CharactersHaibo Jin, Andy Zhou, Joe D. Menke, Haohan WangNeurIPS 2024 · 被引用 55 次
相关 Paper
- Reasoning Models Sometimes Output Illegible Chains of ThoughtArun JoseNeurIPS 2025 · 被引用 11 次
- Large language models can learn and generalize steganographic chain-of-thought under process supervisionRobert MC Carthy, Joey Skaf, Luis Ibañez-Lissen, Vasil Georgiev 等NeurIPS 2025 · 被引用 29 次
- Reasoning Models Struggle to Control their Chains of ThoughtChen Yueh-Han, Robert McCarthy, Bruce W. Lee, He He 等ICML 2026 · 被引用 11 次
- Output Supervision Can Obfuscate the Chain of ThoughtJacob Drori, Luke Marks, Bryce Woodworth, Alex Cloud 等ICLR 2026 · 被引用 10 次
- Early Signs of Steganographic Capabilities in Frontier LLMsArtur Zolkowski, Kei Nishimura-Gasparian, Robert McCarthy, Roland S. Zimmermann 等ICLR 2026 · 被引用 22 次
