Uncertainty Propagation on LLM Agent
Qiwei Zhao, Dong Li, Yanchi Liu, Wei Cheng, Yiyou Sun, Mika Oishi, Takao Osaki, Katsushi Matsuda, Huaxiu Yao, Chen Zhao, Haifeng Chen, Xujiang Zhao
摘要
Large language models (LLMs) integrated into multistep agent systems enable complex decision-making processes across various applications. However, their outputs often lack reliability, making uncertainty estimation crucial. Existing uncertainty estimation methods primarily focus on final-step outputs, which fail to account for cumulative uncertainty over the multistep decision-making process and the dynamic interactions between agents and their environments. To address these limitations, we propose SAUP (Situation Awareness Uncertainty Propagation), a novel framework that propagates uncertainty through each step of an LLM-based agent's reasoning process. SAUP incorporates situational awareness by assigning situational weights to each step's uncertainty during the propagation. Our method, compatible with various one-step uncertainty estimation techniques, provides a comprehensive and accurate uncertainty measure. Extensive experiments on benchmark datasets demonstrate that SAUP significantly outperforms existing state-of-the-art methods, achieving up to 20% improvement in AUROC.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Uncertainty Quantification in LLM Agents: Foundations, Emerging Challenges, and OpportunitiesChangdae Oh, Seongheon Park, To Eun Kim, Jiatong Li 等ACL 2026 · 被引用 8 次
- Agentic Confidence CalibrationJiaxin Zhang, Caiming Xiong, Chien-Sheng WuICML 2026 · 被引用 7 次
- TRACER: Trajectory Risk Aggregation for Critical Episodes in Agentic ReasoningSina Tayebati, Divake Kumar, Nastaran Darabi, Davide Ettori 等ICML 2026 · 被引用 7 次
- Every Response Counts: Quantifying Uncertainty of LLM-based Multi-Agent Systems through Tensor DecompositionTiejin Chen, Huaiyuan Yao, Jia Chen, Evangelos E. Papalexakis 等ACL 2026 · 被引用 3 次
- CoDA-Bench: Can Code Agents Handle Data-Intensive Tasks?Yuxin Zhang, Ju Fan, Meihao Fan, Shaolei Zhang 等ICML 2026 · 被引用 2 次
它引用的顶会 Paper7
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model SocietyGuohao Li, Hasan Hammoud, Hani Itani, Dmitrii Khizbullin 等NeurIPS 2023 · 被引用 1,975 次
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 被引用 1,224 次
- Uncertainty Estimation in Autoregressive Structured PredictionAndrey Malinin, Mark J. F. GalesICLR 2021 · 被引用 439 次
- Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language GenerationLorenz Kuhn, Yarin Gal, Sebastian FarquharICLR 2023 · 被引用 49 次
相关 Paper
- CER: Confidence Enhanced Reasoning in LLMsAli Razghandi, Seyed Mohammad Hadi Hosseini, Mahdieh Soleymani BaghshahACL 2025 · 被引用 11 次
- TokUR: Token-Level Uncertainty Estimation for Large Language Model ReasoningTunyu Zhang, Haizhou Shi, Yibin Wang, Hengyi Wang 等ICLR 2026 · 被引用 19 次
- PPDL: LLM-Based Flows as Probabilistic ProgramsLouis Mandel, Guillaume Baudart, Mandana Vaziri, Martin HirzelICML 2026
- Uncertainty Quantification and Decomposition for LLM-based RecommendationWonbin Kweon, Sanghwan Jang, SeongKu Kang, Hwanjo YuWWW 2025 · 被引用 13 次
- Towards Reliable LLM-based Robots Planning via Combined Uncertainty EstimationShiyuan Yin, Chenjia Bai, Zihao Zhang, Junwei Jin 等NeurIPS 2025
