Cross-model Transferability among Large Language Models on the Platonic Representations of Concepts
Youcheng Huang, Chen Huang, Duanyu Feng, Wenqiang Lei, Jiancheng Lv
摘要
Understanding the inner workings of Large Language Models (LLMs) is a critical research frontier. Prior work has shown that a single LLM's concept representations can be captured as steering vectors (SVs), enabling the control of LLM behavior (e.g., towards generating harmful content). This paper takes a novel approach by exploring the intricate relationships between representations of concepts across different LLMs, drawing an intriguing parallel to the Plato's Allegory of the Cave. In particular, we introduce a linear transformation method to bridge these representations and present three key findings: 1) The representations of a same concept in different LLMs can be effectively aligned using simple linear transformations, enabling efficient cross-model transfer and behavioral control via SVs. 2) This linear transformation generalizes across multiple concepts, facilitating alignment and control of SVs representing different concepts across LLMs. 3) A weakto-strong transferability exists between LLMs, whereby SVs extracted from smaller LLMs can effectively control behaviors of larger LLMs. 1 * Corresponding author. 1 We will release our code at https://github.com/ HamLaertes/Cross_Model_Trans .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust RefusalMantas Mazeika, Long Phan, Xuwang Yin, Andy Zou 等ICML 2024 · 被引用 1,031 次
- The Linear Representation Hypothesis and the Geometry of Large Language ModelsKiho Park, Yo Joong Choe, Victor VeitchICML 2024 · 被引用 461 次
- Fuzz4All: Universal Fuzzing with Large Language ModelsChunqiu Steven Xia, Matteo Paltenghi, Jia Le Tian, Michael Pradel 等ICSE 2024 · 被引用 155 次
- Concept Algebra for (Score-Based) Text-Controlled Generative ModelsZihao Wang, Lin Gui, Jeffrey Negrea, Victor VeitchNeurIPS 2023 · 被引用 77 次
- On the Origins of Linear Representations in Large Language ModelsYibo Jiang, Goutham Rajendran, Pradeep Kumar Ravikumar, Bryon Aragam 等ICML 2024 · 被引用 68 次
相关 Paper
- Analysing the Generalisation and Reliability of Steering VectorsDaniel Tan, David Chanin, Aengus Lynch, Brooks Paige 等NeurIPS 2024
- SAE-SSV: Supervised Steering in Sparse Representation Spaces for Reliable Control of Language ModelsZirui He, Mingyu Jin, Bo Shen, Ali Payani 等EMNLP 2025 · 被引用 1 次
- Separating Tongue from Thought: Activation Patching Reveals Language-Agnostic Concept Representations in TransformersClément Dumas, Chris Wendler, Veniamin Veselovsky, Giovanni Monea 等ACL 2025
- Concept Heterogeneity-aware Representation SteeringLaziz Abdullaev, Noelle Y. L. Wong, Ryan Lee, Shiqi Jiang 等ICML 2026
- Textual Steering Vectors Can Improve Visual Understanding in Multimodal Large Language ModelsWoody Haosheng Gan, Deqing Fu, Julian Asilis, Ollie Liu 等ACL 2026 · 被引用 6 次
