Language Models are Injective and Hence Invertible
Giorgos Nikolaou, Tommaso Mencattini, Donato Crisostomi, Andrea Santilli, Yannis Panagakis, Emanuele Rodolà
摘要
Transformer components such as non-linear activations and normalization are inherently non-injective, suggesting that different inputs could map to the same output and prevent exact recovery of the input from a model's representations. In this paper, we challenge this view. First, we prove mathematically that transformer language models mapping discrete input sequences to their corresponding sequence of continuous representations are injective and therefore lossless, a property established at initialization and preserved during training. Second, we confirm this result empirically through billions of collision tests on six state-of-the-art language models, and observe no collisions. Third, we operationalize injectivity: we introduce SIPIT, the first algorithm that provably and efficiently reconstructs the exact input text from hidden activations, establishing linear-time guarantees and demonstrating exact invertibility in practice. Overall, our work establishes injectivity as a fundamental and exploitable property of language models, with direct implications for transparency, interpretability, and safe deployment. * Equal contribution; author order settled via Mario Kart.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Privacy Control in Conversational LLM Platforms: A Walkthrough StudyZhuoyang Li, Yanlai Wu, Yao Li, Xinning Gui 等CHI 2026 · 被引用 1 次
- Operationalising the Superficial Alignment Hypothesis via Task ComplexityTomás Vergara Browne, Darshan Patil, Ivan Titov, Siva Reddy 等ICML 2026
- PragLocker: Protecting Agent Intellectual Property in Untrusted Deployments via Non-Portable PromptsQinfeng Li, Yuntai Bao, Jianghui Hu, Wenqi Zhang 等ICML 2026
- HijackKV: New Threat in Position-Independent KV Cache ReuseYichi Zhang, Zhiqi Wang, Huan Zhang, Yuchen YangUSENIX Security 2026
它引用的顶会 Paper10
- AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated PromptsTaylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace 等EMNLP 2020 · 被引用 1,162 次
- Attention is not all you need: pure attention loses rank doubly exponentially with depthYihe Dong, Jean-Baptiste Cordonnier, Andreas LoukasICML 2021 · 被引用 522 次
- Are Transformers universal approximators of sequence-to-sequence functions?Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank J. Reddi 等ICLR 2020 · 被引用 481 次
- Hard Prompts Made Easy: Gradient-Based Discrete Optimization for Prompt Tuning and DiscoveryYuxin Wen, Neel Jain, John Kirchenbauer, Micah Goldblum 等NeurIPS 2023 · 被引用 454 次
- RLPrompt: Optimizing Discrete Text Prompts with Reinforcement LearningMingkai Deng, Jianyu Wang, Cheng-Ping Hsieh, Yihan Wang 等EMNLP 2022 · 被引用 141 次
相关 Paper
- Better Language Model Inversion by Compactly Representing Next-Token DistributionsMurtaza Nazir, Matthew Finlayson, John X. Morris, Xiang Ren 等NeurIPS 2025 · 被引用 13 次
- Transformers in Pseudo-Random Number Generation: A Dual Perspective on Theory and PracticeRan Li, Lingshu ZengAAAI 2026
- LLM Self-Recognition: Steering and Retrieving Activation SignaturesThibaud Ardoin, Jonas Schäfer, Gerhard WunderICML 2026
- Language Models Are Implicitly ContinuousSamuele Marro, Davide Evangelista, Xuanqiang Angelo Huang, Emanuele La Malfa 等ICLR 2025
- Token-wise Decomposition of Autoregressive Language Model Hidden States for Analyzing Model PredictionsByung-Doh Oh, William SchulerACL 2023 · 被引用 1 次
