Latent Concept Disentanglement in Transformer-based Language Models
Guanzhe Hong, Bhavya Vasudeva, Vatsal Sharan, Cyrus Rashtchian, Prabhakar Raghavan, Rina Panigrahy
Abstract
When large language models (LLMs) use in-context learning (ICL) to solve a new task, they must infer latent concepts from demonstration examples. This raises the question of whether and how transformers represent latent structures as part of their computation. Our work experiments with several controlled tasks, studying this question using mechanistic interpretability. First, we show that in transitive reasoning tasks with a latent, discrete concept, the model successfully identifies the latent concept and does step-by-step concept composition. This builds upon prior work that analyzes single-step reasoning. Then, we consider tasks parameterized by a latent numerical concept. We discover low-dimensional subspaces in the model's representation space, where the geometry cleanly reflects the underlying parameterization. Overall, we show that small and large models can indeed disentangle and utilize latent concepts that they learn in-context from a handful of abbreviated demonstrations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4024c0a2-58ea-47ec-acd0-e166dde60e17Cited by top-tier papers2
- A Implies B: Circuit Analysis in LLMs for Propositional Logical ReasoningGuanzhe Hong, Nishanth Dikkala, Enming Luo, Cyrus Rashtchian et al.NeurIPS 2025 · 17 citations
- Are Transformers Able to Reason by Connecting Separated Knowledge in Training Data?Yutong Yin, Zhaoran WangICLR 2025
Builds on33
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Implicit In-context LearningZhuowei Li, Zihao Xu, Ligong Han, Yunhe Gao et al.ICLR 2025 · 1,989 citations
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka et al.NeurIPS 2024 · 1,166 citations
- What Can Transformers Learn In-Context? A Case Study of Simple Function ClassesShivam Garg, Dimitris Tsipras, Percy Liang, Gregory ValiantNeurIPS 2022 · 883 citations
Related papers
- Task Generalization with Autoregressive Compositional Structure: Can Learning from D Tasks Generalize to DT Tasks?Amirhesam Abedsoltan, Huaqing Zhang, Kaiyue Wen, Hongzhou Lin et al.ICML 2025
- Unlocking the Black Box of Latent Reasoning: An Interpretability-Guided Approach to InterventionShuochen Chang, Tong Bai, Xiaofeng Zhang, Qianli Ma et al.ACL 2026 · 1 citation
- How Do Transformers Learn In-Context Beyond Simple Functions? A Case Study on Learning with RepresentationsTianyu Guo, Wei Hu, Song Mei, Huan Wang et al.ICLR 2024 · 80 citations
- Chain-of-Thought Provably Enables Learning the (Otherwise) UnlearnableChenxiao Yang, Zhiyuan Li, David WipfICLR 2025
- How do Transformers Learn Implicit Reasoning?Jiaran Ye, Zijun Yao, Zhidian Huang, Liangming Pan et al.NeurIPS 2025 · 17 citations
