Preemptive Detection and Correction of Misaligned Actions in LLM Agents
Haishuo Fang, Xiaodan Zhu, Iryna Gurevych
摘要
Deploying LLM-based agents in real-life applications often faces a critical challenge: the misalignment between agents' behavior and user intent.Such misalignment may lead agents to unintentionally execute some critical actions that carry negative outcomes (e.g., accidentally triggering a 'buy-now' in web shopping), resulting in undesirable or even irreversible consequences.Although addressing these issues is crucial, the preemptive detection and correction of misaligned actions remains relatively underexplored.To fill this gap, we introduce InferAct, a novel approach that leverages the belief reasoning ability of LLMs, grounded in Theory-of-Mind, to detect misaligned actions before execution.Once the misalignment is detected, InferAct alerts users for timely correction, preventing adverse outcomes and enhancing the reliability of LLM agents' decision-making processes.Experiments on three widely used tasks demonstrate InferAct achieves up to 20% improvements on Marco-F1 against baselines in misaligned action detection.An in-depth evaluation of misalignment correction further highlights InferAct's effectiveness in improving agent alignment. 1I want to find white blackout shades that are 66 inches in width and 66 inches in height.they need to be easy to install, and price lower than 90.00 dollars. Thought: I think I should use the search bar to look for white blackout shades that are 66x66 Action: search[white blackout shades 66x66 inches easy to install price < 90.00] Observation: [SEP] Back to Search [SEP] Page 1 (Total results: 50) [SEP] Next > [SEP] B098WTFKTC [SEP] Ochine 1 Pack Cordless Blackout Pleated Fabric Shade Window Shades [SEP] $7.99 [SEP] B0126EWYNI [SEP] Redi Shade Easy Lift Trim-At-Home Cordless ... Thought:The product B09LS7KQMC seems to offer custom cut to size cellular shades, which could potentially meet the size requirement of 66 inches in width and 66 inches in height, and it is also within the price range specified....
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper19
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- WebShop: Towards Scalable Real-World Web Interaction with Grounded Language AgentsShunyu Yao, Howard Chen, John Yang, Karthik NarasimhanNeurIPS 2022 · 被引用 1,477 次
- WebArena: A Realistic Web Environment for Building Autonomous AgentsShuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou 等ICLR 2024 · 被引用 1,197 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
相关 Paper
- Think Twice Before You Act: Enhancing Agent Behavioral Safety with Thought CorrectionChangyue Jiang, Wenqi Zhang, Xudong Pan, Geng Hong 等ICML 2026 · 被引用 13 次
- Executable Code Actions Elicit Better LLM AgentsXingyao Wang, Yangyi Chen, Lifan Yuan, Yizhe Zhang 等ICML 2024 · 被引用 436 次
- Can LLM Agents Simulate Multi-Turn Human Behavior? Evidence from Real Online Customer Behavior DataYuxuan Lu, Jing Huang, Yan Han, Bingsheng Yao 等ACL 2026 · 被引用 17 次
- Expectation Alignment: Handling Reward Misspecification in the Presence of Expectation MismatchMalek Mechergui, Sarath SreedharanNeurIPS 2024 · 被引用 4 次
- AgentOccam: A Simple Yet Strong Baseline for LLM-Based Web AgentsKe Yang, Yao Liu, Sapana Chaudhary, Rasool Fakoor 等ICLR 2025 · 被引用 3 次
