Demystifying optimized prompts in language models
Rimon Melamed, Lucas H. McCabe, H. Howie Huang
摘要
Modern language models (LMs) are not robust to out-of-distribution inputs. Machine generated (``optimized'') prompts can be used to modulate LM outputs and induce specific behaviors while appearing completely uninterpretable. In this work, we investigate the composition of optimized prompts, as well as the mechanisms by which LMs parse and build predictions from optimized prompts. We find that optimized prompts primarily consist of punctuation and noun tokens which are more rare in the training data. Internally, optimized prompts are clearly distinguishable from natural language counterparts based on sparse subsets of the model's activations. Across various families of instruction-tuned models, optimized prompts follow a similar path in how their representations form through the network.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated PromptsTaylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace 等EMNLP 2020 · 被引用 1,162 次
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart 等ICLR 2024 · 被引用 1,072 次
- Hard Prompts Made Easy: Gradient-Based Discrete Optimization for Prompt Tuning and DiscoveryYuxin Wen, Neel Jain, John Kirchenbauer, Micah Goldblum 等NeurIPS 2023 · 被引用 454 次
相关 Paper
- At the Edge of Understanding: Sparse Autoencoders Trace The Limits of Transformer GeneralizationPraneet Suresh, Jack Stanley, Sonia Joseph, Luca Scimeca 等ICML 2026
- Distribution Prompting: Understanding the Expressivity of Language Models Through the Next-Token Distributions They Can ProduceHaojin Wang, Zining Zhu, Freda ShiEMNLP 2025
- On Prompt-Driven Safeguarding for Large Language ModelsChujie Zheng, Fan Yin, Hao Zhou, Fandong Meng 等ICML 2024 · 被引用 116 次
- Evaluating the Zero-shot Robustness of Instruction-tuned Language ModelsJiuding Sun, Chantal Shaib, Byron C. WallaceICLR 2024 · 被引用 75 次
- ASIDE: Architectural Separation of Instructions and Data in Language ModelsEgor Zverev, Evgenii Kortukov, Alexander Panfilov, Alexandra Volkova 等ICLR 2026 · 被引用 28 次
