VeriPlan: Integrating Formal Verification and LLMs into End-User Planning
Christine P. Lee, David Porfirio, Xinyu Jessica Wang, Kevin Chenkai Zhao, Bilge Mutlu
Abstract
Automated planning is traditionally the domain of experts, utilized in fields like manufacturing and healthcare with the aid of expert planning tools. Recent advancements in LLMs have made planning more accessible to everyday users due to their potential to assist users with complex planning tasks. However, LLMs face several application challenges within end-user planning, including consistency, accuracy, and user trust issues. This paper introduces VeriPlan, a system that applies formal verification techniques, specifically model checking, to enhance the reliability and flexibility of LLMs for end-user planning. In addition to the LLM planner, VeriPlan includes three additional core features -- a rule translator, flexibility sliders, and a model checker -- that engage users in the verification process. Through a user study (n=12), we evaluate VeriPlan, demonstrating improvements in the perceived quality, usability, and user satisfaction of LLMs. Our work shows the effective integration of formal verification and user-control features with LLMs for end-user planning tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2786b836-5938-46c7-8dc2-063f069d1f23Cited by top-tier papers3
- ACE: A Security Architecture for LLM-Integrated App SystemsEvan Li, Tushin Mallick, Evan Rose, William K. Robertson et al.NDSS 2026 · 57 citations
- Grounding Generative Planners in Verifiable Logic: A Hybrid Architecture for Trustworthy Embodied AIFeiyu Wu, Xu Zheng, Yue Qu, Zhuocheng Wang et al.ICLR 2026 · 4 citations
- "Do I Trust the AI?" Towards Trustworthy AI-Assisted Diagnosis: Understanding User Perception in LLM-Supported Clinical ReasoningYuansong Xu, Yichao Zhu, Haokai Wang, Yuchen Wu et al.CHI 2026 · 1 citation
Builds on29
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 1,792 citations
- Teaching Large Language Models to Self-DebugXinyun Chen, Maxwell Lin, Nathanael Schärli, Denny ZhouICLR 2024 · 1,085 citations
- Why Johnny Can't Prompt: How Non-AI Experts Try (and Fail) to Design LLM PromptsJ. D. Zamfirescu-Pereira, Richmond Y. Wong, Bjoern Hartmann, Qian YangCHI 2023 · 892 citations
- ChatEval: Towards Better LLM-based Evaluators through Multi-Agent DebateChi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu et al.ICLR 2024 · 871 citations
Related papers
- PAT-Agent: Autoformalization for Model CheckingXinyue Zuo, Yifan Zhang, Hongshu Wang, Yufan Cai et al.ASE 2025 · 1 citation
- VeriSimpl: Robust Optimization Modeling from Natural Language using Simplification-based VerificationSumaya Abdul Rahman, Seckhen Cuellar, Ghani Raissov, Mohammad RazaICML 2026 · 1 citation
- Lemur: Integrating Large Language Models in Automated Program VerificationHaoze Wu, Clark W. Barrett, Nina NarodytskaICLR 2024 · 67 citations
- Can LLMs Fix Issues with Reasoning Models? Towards More Likely Models for AI PlanningTurgay Caglar, Sirine Belhaj, Tathagata Chakraborti, Michael Katz et al.AAAI 2024 · 11 citations
- VeriSafe Agent: Safeguarding Mobile GUI Agent via Logic-based Action VerificationJungjae Lee, Dongjae Lee, Chihun Choi, Youngmin Im et al.MobiCom 2025 · 3 citations
