DLM-Scope: Mechanistic Interpretability of Diffusion Language Models via Sparse Autoencoders
Xu Wang, Bingqing Jiang, Yu Wan, Baosong Yang, Lingpeng Kong, Difan Zou
Abstract
Sparse autoencoders (SAEs) have become a standard tool for mechanistic interpretability in autoregressive large language models (LLMs), enabling researchers to extract sparse, human-interpretable features and intervene on model behavior. Recently, as diffusion language models (DLMs) have become an increasingly powerful and promising alternative to autoregressive LLMs, it is essential to develop tailored mechanistic interpretability tools for this emerging class of models. In this work, we present DLM-Scope , the first SAE-based interpretability framework for DLMs, and demonstrate that trained Top-K SAEs can faithfully extract sparse, interpretable features. Notably, we find that inserting SAEs affects DLMs differently from autoregressive LLMs: while SAE insertion in LLMs typically incurs a loss penalty, in DLMs it can reduce cross-entropy loss when applied to early layers, a phenomenon absent or markedly weaker in LLMs. Additionally, SAE features in DLMs enable more effective diffusion-time interventions, often outperforming LLM steering. Moreover, we pioneer new SAE-based research directions for DLMs: we show that SAEs provide useful signals for DLM decoding order, and that SAE features remain stable during DLM post-training. Overall, our work establishes a foundation for mechanistic interpretability in DLMs and highlights the potential of applying SAEs to DLM-related tasks and algorithms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 67978891-451a-421f-b17f-ad43386b8307Builds on29
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Structured Denoising Diffusion Models in Discrete State-SpacesJacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow et al.NeurIPS 2021 · 2,256 citations
- Diffusion-LM Improves Controllable Text GenerationXiang Lisa Li, John Thickstun, Ishaan Gulrajani, Percy Liang et al.NeurIPS 2022 · 1,546 citations
- Large Language Diffusion ModelsShen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang et al.NeurIPS 2025 · 949 citations
- AR-Diffusion: Auto-Regressive Diffusion Model for Text GenerationTong Wu, Zhihao Fan, Xiao Liu, Hai-Tao Zheng et al.NeurIPS 2023 · 170 citations
Related papers
- Route Sparse Autoencoder to Interpret Large Language ModelsWei Shi, Sihang Li, Tao Liang, Mingyang Wan et al.EMNLP 2025 · 1 citation
- Automatically Interpreting Millions of Features in Large Language ModelsGonçalo Paulo, Alex Mallen, Caden Juang, Nora BelroseICML 2025
- Does Higher Interpretability Imply Better Utility? A Pairwise Analysis on Sparse AutoencodersXu Wang, Yan Hu, Benyou Wang, Difan ZouICLR 2026 · 9 citations
- Sparse Autoencoders Do Not Find Canonical Units of AnalysisPatrick Leask, Bart Bussmann, Michael T. Pearce, Joseph Isaac Bloom et al.ICLR 2025
- Evaluating SAE interpretability without generating explanationsGonçalo Paulo, Nora BelroseICLR 2026 · 2 citations
