Gaussian Process Probes (GPP) for Uncertainty-Aware Probing
Zi Wang, Alexander Ku, Jason Baldridge, Tom Griffiths, Been Kim
Abstract
Understanding which concepts models can and cannot represent has been fundamental to many tasks: from effective and responsible use of models to detecting out of distribution data. We introduce Gaussian process probes (GPP), a unified and simple framework for probing and measuring uncertainty about concepts represented by models. As a Bayesian extension of linear probing methods, GPP asks what kind of distribution over classifiers (of concepts) is induced by the model. This distribution can be used to measure both what the model represents and how confident the probe is about what the model represents. GPP can be applied to any pre-trained model with vector representations of inputs (e.g., activations). It does not require access to training data, gradients, or the architecture. We validate GPP on datasets containing both synthetic and real images. Our experiments show it can (1) probe a model's representations of concepts even with a very small number of examples, (2) accurately measure both epistemic uncertainty (how confident the probe is) and aleatory uncertainty (how fuzzy the concepts are to the model), and (3) detect out of distribution data using those uncertainty measures as well as classic methods do. By using Gaussian processes to expand what probing can offer, GPP provides a data-efficient, versatile and uncertainty-aware tool for understanding and evaluating the capabilities of machine learning models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 943ea5a4-5d45-4356-92e0-9b942310884dCited by top-tier papers3
- Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language ModelsAsma Ghandeharioun, Avi Caciularu, Adam Pearce, Lucas Dixon et al.ICML 2024 · 197 citations
- InversionView: A General-Purpose Method for Reading Information from Neural ActivationsXinting Huang, Madhur Panwar, Navin Goyal, Michael HahnNeurIPS 2024 · 10 citations
- Locate-then-edit for Multi-hop Factual Recall under Knowledge EditingZhuoran Zhang, Yongxiang Li, Zijian Kan, Keyuan Cheng et al.ICML 2025
Builds on7
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Out-of-Distribution Detection with Deep Nearest NeighborsYiyou Sun, Yifei Ming, Xiaojin Zhu, Yixuan LiICML 2022 · 789 citations
- A Theory of Usable Information under Computational ConstraintsYilun Xu, Shengjia Zhao, Jiaming Song, Russell Stewart et al.ICLR 2020 · 211 citations
- Perturbed Masking: Parameter-free Probing for Analyzing and Interpreting BERTZhiyong Wu, Yun Chen, Ben Kao, Qun LiuACL 2020 · 158 citations
- Probing for the Usage of Grammatical NumberKarim Lasri, Tiago Pimentel, Alessandro Lenci, Thierry Poibeau et al.ACL 2022 · 72 citations
Related papers
- What Would Gauss Say About Representations? Probing Pretrained Image Models using Synthetic Gaussian BenchmarksChing-Yun Ko, Pin-Yu Chen, Payel Das, Jeet Mohapatra et al.ICML 2024
- Beyond Single Concept Vector: Modeling Concept Subspace in LLMs with Gaussian DistributionHaiyan Zhao, Heng Zhao, Bo Shen, Ali Payani et al.ICLR 2025 · 2 citations
- Epistemic Uncertainty Quantification for Pretrained Neural NetworksHanjing Wang, Qiang JiCVPR 2024 · 5 citations
- Epistemic Uncertainty for Generated Image DetectionJun Nie, Yonggang Zhang, Tongliang Liu, Yiu-ming Cheung et al.NeurIPS 2025 · 3 citations
- Distinguishing the Knowable from the Unknowable with Language ModelsGustaf Ahdritz, Tian Qin, Nikhil Vyas, Boaz Barak et al.ICML 2024 · 44 citations
