Path Drift in Large Reasoning Models: How First-Person Commitments Override Safety
Yuyi Huang, Runzhe Zhan, Lidia S. Chao, Ailin Tao, Derek F. Wong
摘要
As large reasoning models are increasingly deployed for complex reasoning tasks, Chainof-Thought prompting has emerged as a key paradigm for structured inference. Despite early-stage safeguards enabled by alignment techniques such as RLHF, we identify a previously underexplored vulnerability: reasoning trajectories in LRMs can drift from aligned paths, resulting in content that violates safety constraints. We term this phenomenon Path Drift. Through empirical analysis, we uncover three behavioral triggers of Path Drift: (1) firstperson commitments that induce goal-driven reasoning that delays refusal signals; (2) ethical evaporation, where surface-level disclaimers bypass alignment checkpoints; and (3) condition chain escalation, where layered cues progressively steer models toward unsafe completions. Building on these insights, we introduce a three-stage Path Drift Induction Framework comprising cognitive load amplification, selfrole priming, and condition chain hijacking. Each stage independently reduces refusal rates, while their combination further compounds the effect. To mitigate these risks, we propose a path-level defense strategy incorporating role attribution correction and metacognitive reflection. Our findings highlight the need for trajectory-level alignment oversight in longform reasoning beyond token-level alignment. Warning: This paper contains jailbreak contents that can be offensive in nature.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper6
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Least-to-Most Prompting Enables Complex Reasoning in Large Language ModelsDenny Zhou, Nathanael Schärli, Le Hou, Jason Wei 等ICLR 2023 · 被引用 318 次
- When More is Less: Understanding Chain-of-Thought Length in LLMsYuyang Wu, Yifei Wang, Ziyu Ye, Tianqi Du 等ICLR 2026 · 被引用 225 次
- Jailbreaking LLMs with Arabic Transliteration and ArabiziMansour Al Ghanim, Saleh Almohaimeed, Mengxin Zheng, Yan Solihin 等EMNLP 2024 · 被引用 3 次
相关 Paper
- AdvChain: Adversarial Chain-of-Thought Tuning for Robust Safety Alignment of Large Reasoning ModelsZihao Zhu, Xinyu Wu, Gehan Hu, Siwei Lyu 等ICLR 2026 · 被引用 6 次
- SAFEPATH: Preventing Harmful Reasoning in Chain-of-Thought via Early AlignmentWonje Jeung, Sangyeon Yoon, Minsuk Kahng, Albert NoNeurIPS 2025 · 被引用 31 次
- Towards Safe Reasoning in Large Reasoning Models via Corrective InterventionYichi Zhang, Yue Ding, Jingwen Yang, Tianwei Luo 等ICLR 2026 · 被引用 13 次
- ReasoningGuard: Safeguarding Large Reasoning Models with Inference-time Safety Aha MomentsYuquan Wang, Mi Zhang, Yining Wang, Geng Hong 等ACL 2026 · 被引用 2 次
- Characterizing and Mitigating Reasoning Drift in Large Language ModelsYufeng Zhang, Xuepeng Wang, Lingxiang Wu, Jinqiao WangICLR 2026
