From Associations to Activations: Comparing Behavioral and Hidden-State Semantic Geometry in LLMs
Louis Schiekiera, Max Zimmer, Christophe Roux, Sebastian Pokutta, Fritz Günther
Abstract
We investigate the extent to which an LLM’s hidden-state geometry can be recovered from its behavior in psycholinguistic experiments. Across eight instruction-tuned transformer models, we run two experimental paradigms---similarity-based forced choice and free association---over a shared 5,000-word vocabulary, collecting 17.5M+ trials to build behavior-based similarity matrices. Using representational similarity analysis, we compare behavioral geometries to layerwise hidden-state similarity and benchmark against FastText, BERT, and cross-model consensus. We find that forced-choice behavior aligns substantially more with hidden-state geometry than free association. In a held-out-words regression, behavioral similarity (especially forced choice) predicts unseen hidden-state similarities beyond lexical baselines and cross-model consensus, indicating that behavior-only measurements retain recoverable information about internal semantic geometry. Finally, we discuss implications for the ability of behavioral tasks to uncover hidden cognitive states.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4f47fd4a-9a7f-4911-a8b0-18911eac57adBuilds on9
- Language Models Represent Space and TimeWes Gurnee, Max TegmarkICLR 2024 · 303 citations
- Stealing part of a production language modelNicholas Carlini, Daniel Paleka, Krishnamurthy Dj Dvijotham, Thomas Steinke et al.ICML 2024 · 157 citations
- Interpreting Pretrained Contextualized Representations via Reductions to Static EmbeddingsRishi Bommasani, Kelly Davis, Claire CardieACL 2020 · 137 citations
- Harnessing the Universal Geometry of EmbeddingsRishi D. Jha, Collin Zhang, Vitaly Shmatikov, John X. MorrisNeurIPS 2025 · 69 citations
- Conceptual structure coheres in human cognition but not in large language modelsSiddharth Suresh, Kushin Mukherjee, Xizheng Yu, Wei-Chun Huang et al.EMNLP 2023 · 7 citations
Related papers
- Geometry of Decision Making in Language ModelsAbhinav Joshi, Divyanshu Bhatt, Ashutosh ModiNeurIPS 2025 · 12 citations
- Do LLMs dream of elephants (when told not to)? Latent concept association and associative memory in transformersYibo Jiang, Goutham Rajendran, Pradeep Ravikumar, Bryon AragamNeurIPS 2024 · 19 citations
- CogTaskonomy: Cognitively Inspired Task Taxonomy Is Beneficial to Transfer Learning in NLPYifei Luo, Minghui Xu, Deyi XiongACL 2022 · 20 citations
- Circles are like Ellipses, or Ellipses are like Circles? Measuring the Degree of Asymmetry of Static and Contextual Word Embeddings and the Implications to Representation LearningWei Zhang, Murray Campbell, Yang Yu, Sadhana KumaravelAAAI 2021
- Model-Behavior Alignment under Flexible Evaluation: When the Best-Fitting Model Isn't the Right OneItamar Avitan, Tal GolanNeurIPS 2025 · 5 citations
