Detecting High-Stakes Interactions with Activation Probes
Alex McKenzie, Urja Pawar, Phil Blandfort, William Bankes, David Krueger, Ekdeep Singh Lubana, Dmitrii Krasheninnikov
摘要
Monitoring is an important aspect of safely deploying Large Language Models (LLMs). This paper examines activation probes for detecting ``high-stakes''interactions -- where the text indicates that the interaction might lead to significant harm -- as a critical, yet underexplored, target for such monitoring. We evaluate several probe architectures trained on synthetic data, and find them to exhibit robust generalization to diverse, out-of-distribution, real-world data. Probes'performance is comparable to that of prompted or finetuned medium-sized LLM monitors, while offering computational savings of six orders-of-magnitude. These savings are enabled by reusing activations of the model that is being monitored. Our experiments also highlight the potential of building resource-aware hierarchical monitoring systems, where probes serve as an efficient initial filter and flag cases for more expensive downstream analysis. We release our novel synthetic dataset and the codebase at https://github.com/arrrlex/models-under-pressure.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Constitutional Classifiers++: Efficient Production-Grade Defenses against Universal JailbreaksHoagy Cunningham, Jerry Wei, Zihan Wang, Andrew Persic 等ICLR 2026 · 被引用 39 次
- Reasoning Theater: Disentangling Model Beliefs from Chain-of-ThoughtSiddharth Boppana, Annabel Ma, Max Loeffler, Raphaël Sarfati 等ICML 2026 · 被引用 30 次
- Beyond Linear Probes: Dynamic Safety Monitoring for Language ModelsJames Oldfield, Philip Torr, Ioannis Patras, Adel Bibi 等ICLR 2026 · 被引用 16 次
- Strategic Dishonesty Can Undermine AI Safety Evaluations of Frontier LLMsAlexander Panfilov, Evgenii Kortukov, Kristina Nikolic, Matthias Bethge 等ICLR 2026 · 被引用 14 次
- Diffusion Probe: Generated Image Result Prediction Using CNN ProbesBukun Huang, Benlei Cui, Zhizeng Ye, Xuemei Dong 等CVPR 2026 · 被引用 13 次
它引用的顶会 Paper17
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart 等ICLR 2024 · 被引用 1,072 次
- Hybrid LLM: Cost-Efficient and Quality-Aware Query RoutingDujian Ding, Ankur Mallick, Chi Wang, Robert Sim 等ICLR 2024 · 被引用 282 次
- Red Teaming Language Models with Language ModelsEthan Perez, Saffron Huang, H. Francis Song, Trevor Cai 等EMNLP 2022 · 被引用 239 次
- A is for Absorption: Studying Feature Splitting and Absorption in Sparse AutoencodersDavid Chanin, James Wilken-Smith, Tomás Dulka, Hardik Bhatnagar 等NeurIPS 2025 · 被引用 168 次
- Secret Collusion among AI Agents: Multi-Agent Deception via SteganographySumeet Ramesh Motwani, Mikhail Baranchuk, Martin Strohmeier, Vijay Bolina 等NeurIPS 2024 · 被引用 140 次
相关 Paper
- Detecting Strategic Deception with Linear ProbesNicholas Goldowsky-Dill, Bilal Chughtai, Stefan Heimersheim, Marius HobbhahnICML 2025
- Beyond Token Probes: Hallucination Detection via Activation Tensors with ACT-ViTGuy Bar-Shalom, Fabrizio Frasca, Yaniv Galron, Yftah Ziser 等NeurIPS 2025 · 被引用 17 次
- Toxicity Detection for FreeZhanhao Hu, Julien Piet, Geng Zhao, Jiantao Jiao 等NeurIPS 2024 · 被引用 20 次
- LatentQA: Teaching LLMs to Decode Activations Into Natural LanguageAlexander Pan, Lijie Chen, Jacob SteinhardtICLR 2026 · 被引用 31 次
- Probing Language Models for Pre-training Data DetectionZhenhua Liu, Tong Zhu, Chuanyuan Tan, Bing Liu 等ACL 2024 · 被引用 4 次
