A BERTology View of LLM Orchestrations: Token- and Layer-Selective Probes for Efficient Single-Pass Classification
Gonzalo Ariel Meyoyan, Luciano Del Corro
Abstract
Production LLM systems often rely on separate models for safety and other classification-heavy steps, increasing latency, VRAM footprint, and operational complexity. We instead reuse computation already paid for by the serving LLM: we train lightweight probes on its hidden states and predict labels in the same forward pass used for generation. We frame classification as representation selection over the full token-layer hidden-state tensor, rather than committing to a fixed token or fixed layer (e.g., first-token logits or final-layer pooling). To implement this, we introduce a two-stage aggregator that (i) summarizes tokens within each layer and (ii) aggregates across layer summaries to form a single representation for classification. We instantiate this template with direct pooling, a 100K-parameter scoring-attention gate, and a downcast multi-head self-attention (MHA) probe with up to 35M trainable parameters. Across safety and sentiment benchmarks our probes improve over logit-only reuse (e.g., MULI) and are competitive with substantially larger task-specific baselines, while preserving near-serving latency and avoiding the VRAM and latency costs of a separate guard-model pipeline. Multi-backbone experiments on dense and mixture-of-experts architectures (Llama-3.2-3B, GPT-OSS-20B, Qwen3-30B-A3B) confirm that these findings generalize beyond a single model family.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bac97bc4-0d22-4bc8-8761-4d16ec9c6b7bBuilds on4
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time TestYuhui Li, Fangyun Wei, Chao Zhang, Hongyang ZhangNeurIPS 2025 · 347 citations
- Toxicity Detection for FreeZhanhao Hu, Julien Piet, Geng Zhao, Jiantao Jiao et al.NeurIPS 2024 · 20 citations
- LLaMAs Have Feelings Too: Unveiling Sentiment and Emotion Representations in LLaMA Models Through ProbingDario Di Palma, Alessandro De Bellis, Giovanni Servedio, Vito Walter Anelli et al.ACL 2025 · 11 citations
Related papers
- Quantifying Large Language Model Attacks Through the Lens of Model CognitionXiuming Liu, Chaoxiang He, Xuanran Yu, Jichen Chai et al.USENIX Security 2026
- LLM Safety From Within: Detecting Harmful Content with Internal RepresentationsDifan Jiao, Yilun Liu, Ye Yuan, Zhenwei Tang et al.ACL 2026 · 3 citations
- Predicting LLM Output Length via Entropy-Guided RepresentationsHuanyi Xie, Yubin Chen, Liangyu Wang, Lijie Hu et al.ICLR 2026 · 12 citations
- Efficient Training-Free Multi-Token Prediction via Embedding-Space ProbingRaghavv Goel, Mukul Gagrani, Mingu Lee, Christopher LottICML 2026
- SwiftKV: Fast Prefill-Optimized Inference with Knowledge-Preserving Model TransformationAurick Qiao, Zhewei Yao, Samyam Rajbhandari, Yuxiong HeEMNLP 2025 · 1 citation
