Beyond Token Probes: Hallucination Detection via Activation Tensors with ACT-ViT
Guy Bar-Shalom, Fabrizio Frasca, Yaniv Galron, Yftah Ziser, Haggai Maron
Abstract
Detecting hallucinations in Large Language Model-generated text is crucial for their safe deployment. While probing classifiers show promise, they operate on isolated layer-token pairs and are LLM-specific, limiting their effectiveness and hindering cross-LLM applications. In this paper, we introduce a novel approach to address these shortcomings. We build on the natural sequential structure of activation data in both axes (layers tokens) and advocate treating full activation tensors akin to images. We design ACT-ViT, a Vision Transformer-inspired model that can be effectively and efficiently applied to activation tensors and supports training on data from multiple LLMs simultaneously. Through comprehensive experiments encompassing diverse LLMs and datasets, we demonstrate that ACT-ViT consistently outperforms traditional probing techniques while remaining extremely efficient for deployment. In particular, we show that our architecture benefits substantially from multi-LLM training, achieves strong zero-shot performance on unseen datasets, and can be transferred effectively to new LLMs through fine-tuning. Full code is available at https://github.com/BarSGuy/ACT-ViT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a3796aa9-69e8-4597-8d69-2d2af33b6b99Cited by top-tier papers2
- A Geometric Analysis of Small-sized Language Model HallucinationsEmanuele Ricco, Elia Onofri, Lorenzo Cima, Stefano Cresci et al.ICML 2026 · 1 citation
- Attention Sinks as Internal Signals for Hallucination Detection in Large Language ModelsJakub Binkowski, Kamil Adamczewski, Tomasz KajdanowiczICML 2026
Builds on23
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister et al.NeurIPS 2023 · 1,549 citations
- SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language ModelsPotsawee Manakul, Adian Liusie, Mark J. F. GalesEMNLP 2023 · 331 citations
- INSIDE: LLMs' Internal States Retain the Power of Hallucination DetectionChao Chen, Kai Liu, Ze Chen, Yi Gu et al.ICLR 2024 · 281 citations
- HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language ModelsJunyi Li, Xiaoxue Cheng, Xin Zhao, Jian-Yun Nie et al.EMNLP 2023 · 224 citations
Related papers
- Training a Generalist Hallucination Detector across Multiple Domains via Adaptive Layer AggregationXinyi Li, Zhen Fang, Yadan Luo, Yongxin Deng et al.KDD 2026
- From Out-of-Distribution Detection to Hallucination Detection: A Geometric ViewLitian Liu, Reza Pourreza, Yubing Jian, Yao Qin et al.ICML 2026 · 1 citation
- Detecting LLM Hallucination Through Layer-wise Information Deficiency: Analysis of Ambiguous Prompts and Unanswerable QuestionsHazel Kim, Tom A. Lamb, Adel Bibi, Philip Torr et al.EMNLP 2025 · 1 citation
- Adaptive Activation Steering: A Tuning-Free LLM Truthfulness Improvement Method for Diverse Hallucinations CategoriesTianlong Wang, Xianfeng Jiao, Yinghao Zhu, Zhongzhi Chen et al.WWW 2025 · 64 citations
- Reducing Hallucinations in Large Vision-Language Models via Latent Space SteeringSheng Liu, Haotian Ye, James ZouICLR 2025
