Self-Regularization with Sparse Autoencoders for Controllable LLM-based Classification
Xuansheng Wu, Wenhao Yu, Xiaoming Zhai, Ninghao Liu
Abstract
Modern text classification methods heavily rely on contextual embeddings from large language models (LLMs). Compared to human-engineered features, these embeddings provide automatic and effective representations for classification model training. However, they also introduce a challenge: we lose the ability to manually remove unintended features, such as sensitive or task-irrelevant features, to guarantee regulatory compliance or improve the generalizability of classification models. This limitation arises because LLM embeddings are opaque and difficult to interpret. In this paper, we propose a novel framework to identify and regularize unintended features in the LLM latent space. Specifically, we first pre-train a sparse autoencoder (SAE) to extract interpretable features from LLM latent spaces. To ensure the SAE can capture task-specific features, we further fine-tune it on task-specific datasets. In training the classification model, we propose a simple and effective regularizer, by minimizing the similarity between the classifier weights and the identified unintended feature, to remove the impact of these unintended features on classification. We evaluate the proposed framework on three real-world tasks, including toxic chat detection, reward modeling, and disease diagnosis. Results show that the proposed self-regularization framework can improve the classifier's generalizability by regularizing those features that are not semantically correlated to the task. This work pioneers controllable text classification on LLM latent spaces by leveraging interpreted features to address generalizability, fairness, and privacy challenges. The code and data are publicly available at https://github.com/JacksonWuxs/Controllable_LLM_Classifier.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 65f3b2ee-ee43-46a3-8dbf-ce84e7b1c82fCited by top-tier papers2
- NExT-Guard: Training-Free Streaming Safeguard without Token-Level LabelsJunfeng Fang, Nachuan Chen, Houcheng Jiang, Dan Zhang et al.ICML 2026 · 4 citations
- Less is Enough: Synthesizing Diverse Data in Feature Space of LLMsZhongzhi Li, Xuansheng Wu, Yijiang Li, Lijie Hu et al.ICML 2026 · 1 citation
Builds on15
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- WizardLM: Empowering Large Pre-Trained Language Models to Follow Complex InstructionsCan Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng et al.ICLR 2024 · 1,206 citations
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart et al.ICLR 2024 · 1,072 citations
- LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation DatasetLianmin Zheng, Wei-Lin Chiang, Ying Sheng, Tianle Li et al.ICLR 2024 · 419 citations
Related papers
- Measuring and Guiding MonosemanticityRuben Härle, Felix Friedrich, Manuel Brack, Björn Deiseroth et al.NeurIPS 2025 · 12 citations
- Sparse Autoencoder Features for Classifications and TransferabilityJack Gallifant, Shan Chen, Kuleen Sasse, Hugo J. W. L. Aerts et al.EMNLP 2025
- Characterizing Large Language Model Geometry Helps Solve Toxicity Detection and GenerationRandall Balestriero, Romain Cosentino, Sarath ShekkizharICML 2024 · 9 citations
- Large Language Models can Become Strong Self-DetoxifiersChing-Yun Ko, Pin-Yu Chen, Payel Das, Youssef Mroueh et al.ICLR 2025
- SAE-SSV: Supervised Steering in Sparse Representation Spaces for Reliable Control of Language ModelsZirui He, Mingyu Jin, Bo Shen, Ali Payani et al.EMNLP 2025 · 1 citation
