When Planning Fails Despite Correct Execution: On Epistemic Calibration for LLM-Based Multi-Agent Systems
Zehao Wang, shilong jin, Zhao Cao, Lanjun Wang
Abstract
LLM-based multi-agent systems can fail even when planned actions are executed correctly because agents may misjudge their knowledge when evaluating plan feasibility, a phenomenon we term epistemic miscalibration in planning. Unlike execution errors, epistemic miscalibration is latent during planning, as generated plans can remain self-consistent and executable without observable errors; the miscalibration is also dynamic, as new information can alter feasibility assessments, potentially obscuring past miscalibration signals and causing them to recur over time. To address this, we propose the Epistemic Planning Calibration Agentic Workflow (EPC-AW), which assesses whether plans remain supported under varying information conditions rather than directly verifying feasibility. EPC-AW employs Information-consistency-based Plan Selection, selecting plans whose evaluations are stable across agents, together with Consistency-guided Epistemic State Refinement to adapt calibration over time by leveraging past discrepancies to guide future planning. Experiments show that EPC-AW improves system-level success by an average of 9.75%. Code is available in the public repository (https://github.com/wzhSteve/EPC-AW).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d78b31d4-2311-42c2-a81e-0dbc05d47a8cBuilds on6
- GAIA: a benchmark for General AI AssistantsGrégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun et al.ICLR 2024 · 716 citations
- In-the-Flow Agentic System Optimization for Effective Planning and Tool UseZhuofeng Li, Haoxiang Zhang, Seungju Han, Sheng Liu et al.ICLR 2026 · 65 citations
- Interactive Debugging and Steering of Multi-Agent AI SystemsWill Epperson, Gagan Bansal, Victor C. Dibia, Adam Fourney et al.CHI 2025 · 33 citations
- Creativity in LLM-based Multi-Agent Systems: A SurveyYi-Cheng Lin, Kang-Chieh Chen, Zhe-Yan Li, Tzu-Heng Wu et al.EMNLP 2025 · 3 citations
- Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent SystemsShaokun Zhang, Ming Yin, Jieyu Zhang, Jiale Liu et al.ICML 2025
Related papers
- The Confidence Dichotomy: Analyzing and Mitigating Miscalibration in Tool-Use AgentsWeihao Xuan, Qingcheng Zeng, Heli Qi, Yunze Xiao et al.ACL 2026 · 4 citations
- Towards Epistemic-Doxastic Planning with Observation and RevisionThorsten Engesser, Andreas Herzig, Elise PerrotinAAAI 2024 · 3 citations
- Align While Search: Belief-Guided Exploratory Inference for World-Grounded Embodied AgentsSeohui Bae, Jeonghye Kim, Youngchul Sung, Woohyung LimCVPR 2026 · 1 citation
- Can Dependencies Induced by LLM-Agent Workflows Be Trusted?Yu Yao, Yiliao Song, Yian Xie, Mengdan Fan et al.NeurIPS 2025 · 4 citations
- MetaFaith: Faithful Natural Language Uncertainty Expression in LLMsGabrielle Kaili-May Liu, Gal Yona, Avi Caciularu, Idan Szpektor et al.EMNLP 2025
