Controlling Language and Diffusion Models by Transporting Activations
Pau Rodríguez, Arno Blaas, Michal Klein, Luca Zappella, Nicholas Apostoloff, Marco Cuturi, Xavier Suau
摘要
The increasing capabilities of large generative models and their ever more widespread deployment have raised concerns about their reliability, safety, and potential misuse. To address these issues, recent works have proposed to control model generation by steering model activations in order to effectively induce or prevent the emergence of concepts or behaviors in the generated output. In this paper we introduce Activation Transport (ACT), a general framework to steer activations guided by optimal transport theory that generalizes many previous activation-steering works. ACT is modality-agnostic and provides fine-grained control over the model behavior with negligible computational overhead, while minimally impacting model abilities. We experimentally show the effectiveness and versatility of our approach by addressing key challenges in large language models (LLMs) and text-to-image diffusion models (T2Is). For LLMs, we show that ACT can effectively mitigate toxicity, induce arbitrary concepts, and increase their truthfulness. In T2Is, we show how ACT enables fine-grained style control and concept negation. Once upon a time, there was an old man who lived in the forest. He had no family and he spent his days alone collecting mushrooms for food to survive on. 0.5 Once upon a time, there was an amazing woman named Sarah. She had the most beautiful smile and kindest heart you could ever imagine! Sarah loved to play soccer with her friends on Saturday mornings at 9am sharp every week. 1 Once upon a time, the only way to watch football was on TV. The game of soccer had been played in England since 1863 and by the early twentieth century it became one of Britain's most popular sports. * Equal contribution. Published as a conference paper at ICLR 2025 M.6 STYLE PROMPTS Table 15: List of tags generated with Llama-8B-instruct (right) to induce different styles (left). Anime anime style, large expressive eyes, stylized hair, bold outlines, simplified colors, dynamic perspective, exaggerated features, angular shapes, chibis, manga inspired, emotive facial expressions, action sequences, speed lines, cell shading, graphic backgrounds, vibrant palettes Art nouveau Art Nouveau, Alphonse Mucha, Gustav Klimt, flowing lines, organic shapes, floral motifs, geometric patterns, ornamental designs, Jugendstil, Secessionism, symbolism, female figures, gold leaf, intricate details, turn of the century art, early 20th century Impressionism impressionism, Claude Monet, brush strokes, light, color, outdoor scenes, water lilies, haystacks, Rouen Cathedral, reflections, nature, atmospheric, vibrant colors, visible textures, 19th century art, French impressionism Cyberpunk cyberpunk, neon lights, urban jungles, high-tech architecture, augmented reality, AI technology, biopunk, futuristic cities, post-apocalyptic scenes, digital hacking, megacorporations, androids, dystopian societies, cybernetic enhancements, chromed details, glowing neon signs, rain-soaked streets
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- ODESteer: A Unified ODE-Based Steering Framework for LLM AlignmentHongjue Zhao, Haosen Sun, Jiangtao Kong, Xiaochang Li 等ICLR 2026 · 被引用 16 次
- LinEAS: End-to-end Learning of Activation Steering with a Distributional LossPau Rodríguez, Michal Klein, Eleonora Gualdoni, Valentino Maiorca 等NeurIPS 2025 · 被引用 15 次
- Activation Steering with a Feedback ControllerDung Viet Nguyen, Yen Nhi Pham, Hieu M. Vu, Lei Zhang 等ICLR 2026 · 被引用 13 次
- Video Unlearning via Low-Rank Refusal VectorSimone Facchiano, Stefano Saravalle, Matteo Migliarini, Edoardo De Matteis 等ICLR 2026 · 被引用 11 次
- COLD-Steer: Steering Large Language Models via In-Context One-step Learning DynamicsKartik Sharma, Rakshit S. TrivediICLR 2026 · 被引用 8 次
它引用的顶会 Paper28
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
相关 Paper
- A Geometric Information Bottleneck for Activation SteeringToan Doan, Thin Nguyen, Sunil GuptaKDD 2026
- Controlling Large Language Models Through Concept Activation VectorsHanyu Zhang, Xiting Wang, Chengao Li, Xiang Ao 等AAAI 2025 · 被引用 26 次
- Causal-Steer: Disentangled Continuous Style Control without Parallel CorporaQingsong Wang, Chang Yao, Jingyuan ChenICLR 2026
- Breaking Bad Tokens: Detoxification of LLMs Using Sparse AutoencodersAgam Goyal, Vedant Rathi, William Yeh, Yian Wang 等EMNLP 2025 · 被引用 1 次
- Adaptive Activation Steering: A Tuning-Free LLM Truthfulness Improvement Method for Diverse Hallucinations CategoriesTianlong Wang, Xianfeng Jiao, Yinghao Zhu, Zhongzhi Chen 等WWW 2025 · 被引用 64 次
