Preemptive Detection and Correction of Misaligned Actions in LLM Agents
Haishuo Fang, Xiaodan Zhu, Iryna Gurevych
Abstract
Deploying LLM-based agents in real-life applications often faces a critical challenge: the misalignment between agents' behavior and user intent.Such misalignment may lead agents to unintentionally execute some critical actions that carry negative outcomes (e.g., accidentally triggering a 'buy-now' in web shopping), resulting in undesirable or even irreversible consequences.Although addressing these issues is crucial, the preemptive detection and correction of misaligned actions remains relatively underexplored.To fill this gap, we introduce InferAct, a novel approach that leverages the belief reasoning ability of LLMs, grounded in Theory-of-Mind, to detect misaligned actions before execution.Once the misalignment is detected, InferAct alerts users for timely correction, preventing adverse outcomes and enhancing the reliability of LLM agents' decision-making processes.Experiments on three widely used tasks demonstrate InferAct achieves up to 20% improvements on Marco-F1 against baselines in misaligned action detection.An in-depth evaluation of misalignment correction further highlights InferAct's effectiveness in improving agent alignment. 1I want to find white blackout shades that are 66 inches in width and 66 inches in height.they need to be easy to install, and price lower than 90.00 dollars. Thought: I think I should use the search bar to look for white blackout shades that are 66x66 Action: search[white blackout shades 66x66 inches easy to install price < 90.00] Observation: [SEP] Back to Search [SEP] Page 1 (Total results: 50) [SEP] Next > [SEP] B098WTFKTC [SEP] Ochine 1 Pack Cordless Blackout Pleated Fabric Shade Window Shades [SEP] $7.99 [SEP] B0126EWYNI [SEP] Redi Shade Easy Lift Trim-At-Home Cordless ... Thought:The product B09LS7KQMC seems to offer custom cut to size cellular shades, which could potentially meet the size requirement of 66 inches in width and 66 inches in height, and it is also within the price range specified....
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on19
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- WebShop: Towards Scalable Real-World Web Interaction with Grounded Language AgentsShunyu Yao, Howard Chen, John Yang, Karthik NarasimhanNeurIPS 2022 · 1,477 citations
- WebArena: A Realistic Web Environment for Building Autonomous AgentsShuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou et al.ICLR 2024 · 1,197 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
Related papers
- Think Twice Before You Act: Enhancing Agent Behavioral Safety with Thought CorrectionChangyue Jiang, Wenqi Zhang, Xudong Pan, Geng Hong et al.ICML 2026 · 13 citations
- Executable Code Actions Elicit Better LLM AgentsXingyao Wang, Yangyi Chen, Lifan Yuan, Yizhe Zhang et al.ICML 2024 · 436 citations
- Can LLM Agents Simulate Multi-Turn Human Behavior? Evidence from Real Online Customer Behavior DataYuxuan Lu, Jing Huang, Yan Han, Bingsheng Yao et al.ACL 2026 · 17 citations
- Expectation Alignment: Handling Reward Misspecification in the Presence of Expectation MismatchMalek Mechergui, Sarath SreedharanNeurIPS 2024 · 4 citations
- AgentOccam: A Simple Yet Strong Baseline for LLM-Based Web AgentsKe Yang, Yao Liu, Sapana Chaudhary, Rasool Fakoor et al.ICLR 2025 · 3 citations
