Cross-model Transferability among Large Language Models on the Platonic Representations of Concepts
Youcheng Huang, Chen Huang, Duanyu Feng, Wenqiang Lei, Jiancheng Lv
Abstract
Understanding the inner workings of Large Language Models (LLMs) is a critical research frontier. Prior work has shown that a single LLM's concept representations can be captured as steering vectors (SVs), enabling the control of LLM behavior (e.g., towards generating harmful content). This paper takes a novel approach by exploring the intricate relationships between representations of concepts across different LLMs, drawing an intriguing parallel to the Plato's Allegory of the Cave. In particular, we introduce a linear transformation method to bridge these representations and present three key findings: 1) The representations of a same concept in different LLMs can be effectively aligned using simple linear transformations, enabling efficient cross-model transfer and behavioral control via SVs. 2) This linear transformation generalizes across multiple concepts, facilitating alignment and control of SVs representing different concepts across LLMs. 3) A weakto-strong transferability exists between LLMs, whereby SVs extracted from smaller LLMs can effectively control behaviors of larger LLMs. 1 * Corresponding author. 1 We will release our code at https://github.com/ HamLaertes/Cross_Model_Trans .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 934c1e1f-d529-469d-9dac-a2db792f7703Builds on10
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust RefusalMantas Mazeika, Long Phan, Xuwang Yin, Andy Zou et al.ICML 2024 · 1,031 citations
- The Linear Representation Hypothesis and the Geometry of Large Language ModelsKiho Park, Yo Joong Choe, Victor VeitchICML 2024 · 461 citations
- Fuzz4All: Universal Fuzzing with Large Language ModelsChunqiu Steven Xia, Matteo Paltenghi, Jia Le Tian, Michael Pradel et al.ICSE 2024 · 155 citations
- Concept Algebra for (Score-Based) Text-Controlled Generative ModelsZihao Wang, Lin Gui, Jeffrey Negrea, Victor VeitchNeurIPS 2023 · 77 citations
- On the Origins of Linear Representations in Large Language ModelsYibo Jiang, Goutham Rajendran, Pradeep Kumar Ravikumar, Bryon Aragam et al.ICML 2024 · 68 citations
Related papers
- Analysing the Generalisation and Reliability of Steering VectorsDaniel Tan, David Chanin, Aengus Lynch, Brooks Paige et al.NeurIPS 2024
- SAE-SSV: Supervised Steering in Sparse Representation Spaces for Reliable Control of Language ModelsZirui He, Mingyu Jin, Bo Shen, Ali Payani et al.EMNLP 2025 · 1 citation
- Separating Tongue from Thought: Activation Patching Reveals Language-Agnostic Concept Representations in TransformersClément Dumas, Chris Wendler, Veniamin Veselovsky, Giovanni Monea et al.ACL 2025
- Concept Heterogeneity-aware Representation SteeringLaziz Abdullaev, Noelle Y. L. Wong, Ryan Lee, Shiqi Jiang et al.ICML 2026
- Textual Steering Vectors Can Improve Visual Understanding in Multimodal Large Language ModelsWoody Haosheng Gan, Deqing Fu, Julian Asilis, Ollie Liu et al.ACL 2026 · 6 citations
