Estimating the Black-box LLM Uncertainty with Distribution-Aligned Adversarial Distillation
Huizi Cui, Huan Ma, Qilin Wang, Yuhang Gao, Changqing Zhang
Abstract
Large language models (LLMs) have progressed rapidly in complex reasoning and question answering, yet LLM hallucination remains a central bottleneck that hinders practical deployment, especially for commercial black-box LLMs accessible only via APIs. Existing uncertainty quantification methods typically depend on computationally expensive multiple sampling or internal parameters, which prevents real-time estimation and fails to capture information implicit in the blackbox reasoning process. To address this issue, we propose Distribution-Aligned Adversarial Distillation (DisAAD), which introduces a generation-discrimination architecture to guide a lightweight proxy model to learn the highquality regions of the output distribution of the black-box LLM, thus effectively endowing it with the ability to "know whether the blackbox LLM knows or not". Subsequently, we use the proxy model to reproduce the specific responses of the black-box LLM and estimate the corresponding uncertainty based on evidence learning. Extensive experiments have verified the effectiveness and promise of our proposed method, indicating that a proxy model even one that only accounts for 1% of the target LLM's size can achieve reliable uncertainty quantification. Our model and related resources are released at https://github.com/huizi-Cui/ DisAAD .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cfd7e3e0-9a88-4a9b-a505-b656cc3fd4dfBuilds on8
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Uncertainty Estimation in Autoregressive Structured PredictionAndrey Malinin, Mark J. F. GalesICLR 2021 · 439 citations
- LLM-Check: Investigating Detection of Hallucinations in Large Language ModelsGaurang Sriramanan, Siddhant Bharti, Vinu Sankar Sadasivan, Shoumik Saha et al.NeurIPS 2024 · 170 citations
- Large Language Models Must Be Taught to Know What They Don't KnowSanyam Kapoor, Nate Gruver, Manley Roberts, Katie Collins et al.NeurIPS 2024 · 124 citations
- Shifting Attention to Relevance: Towards the Predictive Uncertainty Quantification of Free-Form Large Language ModelsJinhao Duan, Hao Cheng, Shiqi Wang, Alex Zavalny et al.ACL 2024 · 28 citations
Related papers
- Mapping from Meaning: Addressing the Miscalibration of Prompt-Sensitive Language ModelsKyle Cox, Jiawei Xu, Yikun Han, Rong Xu et al.AAAI 2025 · 6 citations
- Kernel Language Entropy: Fine-grained Uncertainty Quantification for LLMs from Semantic SimilaritiesAlexander Nikitin, Jannik Kossen, Yarin Gal, Pekka MarttinenNeurIPS 2024 · 197 citations
- UAlign: Leveraging Uncertainty Estimations for Factuality Alignment on Large Language ModelsBoyang Xue, Fei Mi, Qi Zhu, Hongru Wang et al.ACL 2025 · 10 citations
- Estimating Semantic Alphabet Size for LLM Uncertainty QuantificationLucas H. McCabe, Rimon Melamed, Tom Hartvigsen, H. Howie HuangICLR 2026 · 7 citations
- Making Expert Reasoning Learnable with Self-DistillationEthan Mendes, Jungsoo Park, Alan RitterICML 2026 · 1 citation
