Inspection and Control of Self-Generated-Text Recognition Ability in Llama3-8b-Instruct
Christopher Ackerman, Nina Panickssery
Abstract
It has been reported that LLMs can recognize their own writing. As this has potential implications for AI safety, yet is relatively understudied, we investigate the phenomenon, seeking to establish whether it robustly occurs at the behavioral level, how the observed behavior is achieved, and whether it can be controlled. First, we find that the Llama3-8b-Instruct chat model - but not the base Llama3-8b model - can reliably distinguish its own outputs from those of humans, and present evidence that the chat model is likely using its experience with its own outputs, acquired during post-training, to succeed at the writing recognition task. Second, we identify a vector in the residual stream of the model that is differentially activated when the model makes a correct self-written-text recognition judgment, show that the vector activates in response to information relevant to self-authorship, present evidence that the vector is related to the concept of "self" in the model, and demonstrate that the vector is causally related to the model's ability to perceive and assert self-authorship. Finally, we show that the vector can be used to control both the model's behavior and its perception, steering the model to claim or disclaim authorship by applying the vector to the model's output as it generates it, and steering the model to believe or disbelieve it wrote arbitrary texts by applying the vector to them as the model reads them.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b360180e-8702-4a44-80bb-34689139d1a3Cited by top-tier papers5
- Many-shot JailbreakingCem Anil, Esin Durmus, Nina Panickssery, Mrinank Sharma et al.NeurIPS 2024 · 338 citations
- The Assistant Axis: Situating and Stabilizing the Default Persona of Language ModelsChristina Lu, Jack Gallagher, Jonathan Michala, Kyle Fish et al.ICML 2026 · 64 citations
- Are LLM Evaluators Really Narcissists? Sanity Checking Self-Preference EvaluationsDani Roytburg, Matthew Bozoukov, Matthew Nguyen, Jou Barzdukas et al.ICML 2026
- Telescope: Improving Zero Shot Detection of LLM Generated Content By Measuring Token Repetition ProbabilityChristopher Nassif, Joshua CooperICML 2026
- LLM Self-Recognition: Steering and Retrieving Activation SignaturesThibaud Ardoin, Jonas Schäfer, Gerhard WunderICML 2026
Builds on3
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LLM Evaluators Recognize and Favor Their Own GenerationsArjun Panickssery, Samuel R. Bowman, Shi FengNeurIPS 2024 · 865 citations
- Many-shot JailbreakingCem Anil, Esin Durmus, Nina Panickssery, Mrinank Sharma et al.NeurIPS 2024 · 338 citations
Related papers
- Few-Shot Detection of Machine-Generated Text using Style RepresentationsRafael A. Rivera Soto, Kailin Koch, Aleem Khan, Barry Y. Chen et al.ICLR 2024 · 49 citations
- The HaLLMark Effect: Supporting Provenance and Transparent Use of Large Language Models in Writing with Interactive VisualizationMd. Naimul Hoque, Tasfia Mashiat, Bhavya Ghai, Cecilia D. Shelton et al.CHI 2024 · 48 citations
- A Bayesian Approach to Harnessing the Power of LLMs in Authorship AttributionZhengmian Hu, Tong Zheng, Heng HuangEMNLP 2024 · 1 citation
- Tell me about yourself: LLMs are aware of their learned behaviorsJan Betley, Xuchan Bao, Martín Soto, Anna Sztyber-Betley et al.ICLR 2025 · 2 citations
- Endogenous Resistance to Activation Steering in Language ModelsAlex McKenzie, Keenan Pepper, Stijn Servaes, Martin Leitgab et al.ICML 2026 · 3 citations
