The Shape of Adversarial Influence: Characterizing LLM Latent Spaces with Persistent Homology
Aideen Fay, Inés García-Redondo, Qiquan Wang, Haim Dubossarsky, Anthea Monod
Abstract
Existing interpretability methods for Large Language Models (LLMs) predominantly capture linear directions or isolated features. This overlooks the high-dimensional, relational, and nonlinear geometry of model representations. We apply persistent homology (PH) to characterize how adversarial inputs reshape the geometry and topology of internal representation spaces of LLMs. This phenomenon, especially when considered across operationally different attack modes, remains poorly understood. We analyze six models (3.8B to 70B parameters) under two distinct attacks, indirect prompt injection and backdoor fine-tuning, and show that a consistent topological signature persists throughout. Adversarial inputs induce topological compression, where the latent space becomes structurally simpler, collapsing the latent space from varied, compact, small-scale features into fewer, dominant, large-scale ones. This signature is architecture-agnostic, emerges early in the network, and is highly discriminative across layers. By quantifying the shape of activation point clouds and neuron-level information flow, our framework reveals geometric invariants of representational change that complement existing linear interpretability methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on9
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart et al.ICLR 2024 · 1,072 citations
- The WMDP Benchmark: Measuring and Reducing Malicious Use with UnlearningNathaniel Li, Alexander Pan, Anjali Gopal, Summer Yue et al.ICML 2024 · 390 citations
- On the geometry of generalization and memorization in deep neural networksCory Stephenson, Suchismita Padhy, Abhinav Ganesh, Yue Hui et al.ICLR 2021 · 95 citations
- Stress-Testing Capability Elicitation With Password-Locked ModelsRyan Greenblatt, Fabien Roger, Dmitrii Krasheninnikov, David KruegerNeurIPS 2024 · 47 citations
Related papers
- Characterizing Large Language Model Geometry Helps Solve Toxicity Detection and GenerationRandall Balestriero, Romain Cosentino, Sarath ShekkizharICML 2024 · 9 citations
- Persistent Topological Features in Large Language ModelsYuri Gardinazzi, Karthik Viswanathan, Giada Panerai, Alessio Ansuini et al.ICML 2025
- Towards Scalable Topological RegularizersHiu-Tung Wong, Darrick Lee, Hong YanICLR 2025
- Beyond Mere Token Analysis: A Hypergraph Metric Space Framework for Defending Against Socially Engineered LLM AttacksManohar Kaul, Aditya Saibewar, Sadbhavana BabarICLR 2025
- Experimental Observations of the Topology of Convolutional Neural Network ActivationsEmilie Purvine, Davis Brown, Brett A. Jefferson, Cliff A. Joslyn et al.AAAI 2023 · 21 citations
