Lune

ICML2026Top-tier venue

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring

Guanxu Chen, Jing Shao, Tao Luo, Lijie Hu, Qihao Lin, Dongrui Liu

2026Year
2Citations

Abstract

Large language models (LLMs) are becoming increasingly capable, but the mechanisms of their thinking and decision-making processes remain unclear. Chain-of-thoughts (CoTs) have been commonly utilized to externalize LLMs' thinking, but this strategy fails to accurately reflect LLMs' thinking process. Techniques based on LLMs' hidden representations provide an inner perspective to improve the monitorability of their latent thinking. However, previous methods only try to develop external modules instead of making LLMs themselves easier to monitor. In this paper, we propose a novel method TELLME, improving the transparency of LLMs for monitoring and helping monitors identify unsuitable and sensitive behaviors. Furthermore, we showcase the effectiveness and scalability of TELLME on detoxification tasks, where LLMs achieve consistent improvement among multimodal test sets, distinct architectures, and varying parameter scales. We further analyze how the generalized improvement of TELLME aligns with the optimal transport framework and empirical perspectives.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext e747a5b6-1ca1-46d3-92ec-17270a0d4ba9

Builds on32

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines