SAE as a Crystal Ball: Interpretable Features Predict Cross-domain Transferability of LLMs without Training
Qi Zhang, Yifei Wang, Xiaohan Wang, Jiajun Chai, Guojun Yin, Wei Lin, Yisen Wang
Abstract
In recent years, pre-trained large language models have achieved remarkable success across diverse tasks. Besides the pivotal role of self-supervised pre-training, their effectiveness in downstream applications also depends critically on the post-training process, which adapts models to task-specific data and objectives. However, this process inevitably introduces model shifts that can influence performance in different domains, and how such shifts transfer remains poorly understood. To open up the black box, we propose the SAE-based Transferability Score (STS), a new metric that leverages sparse autoencoders (SAEs) to forecast post-training transferability. Taking supervised fine-tuning as an example, STS identifies shifted dimensions in SAE representations and calculates their correlations with downstream domains, enabling reliable estimation of transferability before fine-tuning. Extensive experiments across multiple models and domains show that STS accurately predicts the transferability of supervised fine-tuning, achieving Pearson correlation coefficients above 0.7 with actual performance changes. Beyond this, we take an initial step toward extending STS to reinforcement learning. We believe that STS can serve as an black interpretable tool for guiding post-training strategies in LLMs. Code is available at https://github.com/PKU-ML/STS.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a7602c25-cedc-4027-94ab-ff454b85f3faBuilds on10
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart et al.ICLR 2024 · 1,072 citations
- Fine-Tuning can Distort Pretrained Features and Underperform Out-of-DistributionAnanya Kumar, Aditi Raghunathan, Robbie Matthew Jones, Tengyu Ma et al.ICLR 2022 · 911 citations
Related papers
- ReTRE: Benchmarking LLM Transfer Robustness with Structure-Preserving VariantsZhongDong Li, Weijie Shi, Yue Cui, Haolun Ma et al.ACL 2026
- SparseRM: A Lightweight Preference Modeling with Sparse AutoencoderDengcan Liu, Jiahao Li, Zheren Fu, Yi Tu et al.AAAI 2026
- DA-Pred: Performance Prediction for Text Summarization under Domain-Shift and Instruct-TuningAnum Afzal, Florian Matthes, Alexander R. FabbriEMNLP 2025
- Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM ReasoningMaggie Ziyu Huan, Yuetai Li, Tuney Zheng, Xiaoyu Xu et al.ICML 2026 · 102 citations
- Calibrating Zero-shot Cross-lingual (Un-)structured PredictionsZhengping Jiang, Anqi Liu, Benjamin Van DurmeEMNLP 2022 · 4 citations
