Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring
Guanxu Chen, Jing Shao, Tao Luo, Lijie Hu, Qihao Lin, Dongrui Liu
Abstract
Large language models (LLMs) are becoming increasingly capable, but the mechanisms of their thinking and decision-making processes remain unclear. Chain-of-thoughts (CoTs) have been commonly utilized to externalize LLMs' thinking, but this strategy fails to accurately reflect LLMs' thinking process. Techniques based on LLMs' hidden representations provide an inner perspective to improve the monitorability of their latent thinking. However, previous methods only try to develop external modules instead of making LLMs themselves easier to monitor. In this paper, we propose a novel method TELLME, improving the transparency of LLMs for monitoring and helping monitors identify unsuitable and sensitive behaviors. Furthermore, we showcase the effectiveness and scalability of TELLME on detoxification tasks, where LLMs achieve consistent improvement among multimodal test sets, distinct architectures, and varying parameter scales. We further analyze how the generalized improvement of TELLME aligns with the optimal transport framework and empirical perspectives.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e747a5b6-1ca1-46d3-92ec-17270a0d4ba9Builds on32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
Related papers
- Chain-of-Thought Tuning: Masked Language Models can also Think Step By Step in Natural Language UnderstandingCaoyun Fan, Jidong Tian, Yitian Li, Wenqing Chen et al.EMNLP 2023 · 3 citations
- How Far Ahead Do LLMs Plan? Uncovering the Latent Horizon in Chain-of-Thought ReasoningLiyan Xu, Mo Yu, Fandong Meng, Jie ZhouICML 2026 · 1 citation
- Output Supervision Can Obfuscate the Chain of ThoughtJacob Drori, Luke Marks, Bryce Woodworth, Alex Cloud et al.ICLR 2026 · 10 citations
- Latent Thinking Optimization: Your Latent Reasoning Language Model Secretly Encodes Reward Signals in Its Latent ThoughtsHanwen Du, Yuxin Dong, Xia NingICLR 2026 · 20 citations
- UniCoTT: A Unified Framework for Structural Chain-of-Thought DistillationXianwei Zhuang, Zhihong Zhu, Zhichang Wang, Xuxin Cheng et al.ICLR 2025
