Unveiling and Manipulating Prompt Influence in Large Language Models
Zijian Feng, Hanzhang Zhou, Zixiao Zhu, Junlang Qian, Kezhi Mao
Abstract
Prompts play a crucial role in guiding the responses of Large Language Models (LLMs). However, the intricate role of individual tokens in prompts, known as input saliency, in shaping the responses remains largely underexplored. Existing saliency methods either misalign with LLM generation objectives or rely heavily on linearity assumptions, leading to potential inaccuracies. To address this, we propose Token Distribution Dynamics (TDD), a blacksimple yet effective approach to unveil and manipulate the role of prompts in generating LLM outputs. TDD leverages the robust interpreting capabilities of the language model head (LM head) to assess input saliency. It projects input tokens into the embedding space and then estimates their significance based on distribution dynamics over the vocabulary. We introduce three TDD variants: forward, backward, and bidirectional, each offering unique insights into token relevance. Extensive experiments reveal that the TDD surpasses state-of-the-art baselines with a big margin in elucidating the causal relationships between prompts and LLM outputs. Beyond mere interpretation, we apply TDD to two prompt manipulation tasks for controlled text generation: zero-shot toxic language suppression and sentiment steering. Empirical results underscore TDD's proficiency in identifying both toxic and sentimental cues in prompts, subsequently mitigating toxicity or modulating sentiment in the generated content.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- UniBias: Unveiling and Mitigating LLM Bias through Internal Attention and FFN ManipulationHanzhang Zhou, Zijian Feng, Zixiao Zhu, Junlang Qian et al.NeurIPS 2024 · 43 citations
- Deception at Scale: Deceptive Designs in 1K LLM-Generated E-Commerce ComponentsZiwei Chen, Jiawen Shen, Luna, Hanyu Zhang et al.CHI 2026 · 3 citations
- Exploring Multidimensional Checkworthiness: Designing AI-assisted Claim Prioritization for Human Fact-checkersHoujiang Liu, Jacek Gwizdka, Matthew LeaseCSCW 2025 · 3 citations
- FreeCtrl: Constructing Control Centers with Feedforward Layers for Learning-Free Controllable Text GenerationZijian Feng, Hanzhang Zhou, Kezhi Mao, Zixiao ZhuACL 2024 · 3 citations
- Hessian-Enhanced Token Attribution (HETA): Interpreting Autoregressive LLMsVishal Pramanik, Maisha Maliha, Nathaniel D. Bastian, Sumit Kumar JhaICLR 2026 · 3 citations
Builds on16
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- Generic Attention-model Explainability for Interpreting Bi-Modal and Encoder-Decoder TransformersHila Chefer, Shir Gur, Lior WolfICCV 2021 · 451 citations
- A Diagnostic Study of Explainability Techniques for Text ClassificationPepa Atanasova, Jakob Grue Simonsen, Christina Lioma, Isabelle AugensteinEMNLP 2020 · 158 citations
- XAI for Transformers: Better Explanations through Conservative PropagationAmeen Ali, Thomas Schnake, Oliver Eberle, Grégoire Montavon et al.ICML 2022 · 144 citations
- Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary SpaceMor Geva, Avi Caciularu, Kevin Ro Wang, Yoav GoldbergEMNLP 2022 · 92 citations
Related papers
- When Smiley Turns Hostile: Interpreting How Emojis Trigger LLMs' ToxicityShiyao Cui, Xijia Feng, Yingkang Wang, Junxiao Yang et al.AAAI 2026
- Toxicity Detection for FreeZhanhao Hu, Julien Piet, Geng Zhao, Jiantao Jiao et al.NeurIPS 2024 · 20 citations
- The Fragile Truth of Saliency: Improving LLM Input Attribution via Attention Bias OptimizationYihua Zhang, Changsheng Wang, Yiwei Chen, Chongyu Fan et al.NeurIPS 2025 · 1 citation
- Whispering Experts: Neural Interventions for Toxicity Mitigation in Language ModelsXavier Suau, Pieter Delobelle, Katherine Metcalf, Armand Joulin et al.ICML 2024 · 31 citations
- Leashing the Inner Demons: Self-Detoxification for Language ModelsCanwen Xu, Zexue He, Zhankui He, Julian J. McAuleyAAAI 2022 · 30 citations
