Action Boundary Blindness: When LLM Agents Cannot Tell Where One Action Ends and Another Begins
Zhangyi Wang, Bingnan Yu, Jiexiang Xu, Zongze Li
摘要
Large language model (LLM) agents excel at multi-step tasks yet frequently exhibit Action Boundary Blindness-the inability to correctly determine action granularity, scope, and completeness. Grounded in Event Segmentation Theory from cognitive science, we formalize three violation types: granularity confusion, scope creep, and boundary ambiguity. We propose four automatic metrics-Action Boundary Score (ABS), Granularity Alignment Rate (GAR), Scope Violation Rate (SVR), and Boundary-Aware Success Rate (BASR)requiring no human annotation. Experiments on 1,655 tasks across six benchmarks (τ -bench, WebArena, ALFWorld, TheAgentCompany, OSWorld) with seven LLMs reveal that: (1) the best model achieves only 0.424 ABS; (2) using a multi-label attribution framework validated by inter-annotator agreement (κ = 0.78), boundary blindness is the primary failure mode in 37.2% of failures (25.8% as sole cause; 55.9% total involvement including contributing factors); (3) under-action dominates at 48.4%; (4) BASR is consistently ∼4 points lower than traditional success rate, exposing "lucky successes." Critically, Explicit Boundary Prompting (EBP) improves ABS by 0.08-0.13 across all models, demonstrating that boundary blindness is better characterized as an elicitation gap rather than a fundamental capability limitation-LLMs possess latent boundary perception not activated by default. This finding has implications for alignment and instruction tuning. We validate metrics through state-based cross-validation and human audit, estimating ∼22% false positive rate from valid alternative paths, with model rankings remaining stable (Spearman ρ = 1.0).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu 等NeurIPS 2023 · 被引用 5,989 次
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- ALFWorld: Aligning Text and Embodied Environments for Interactive LearningMohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk 等ICLR 2021 · 被引用 819 次
- Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language ModelsLei Wang, Wanyu Xu, Yihuai Lan, Zhiqiang Hu 等ACL 2023 · 被引用 249 次
相关 Paper
- CAR-bench: Evaluating the Consistency and Limit-Awareness of LLM Agents under Real-World UncertaintyJohannes Kirmayr, Lukas Stappen, Elisabeth AndréACL 2026 · 被引用 5 次
- Bridging the Gulf of Envisioning: Cognitive Challenges in Prompt Based Interactions with LLMsHariharan Subramonyam, Roy Pea, Christopher Lawrence Pondoc, Maneesh Agrawala 等CHI 2024 · 被引用 137 次
- Scope Delineation Before Localization: A Two-Stage Framework for Enhancing Failure Attribution in Multi-Agent SystemsKai Sun, Wenqiang Li, Bo Dong, Yuxin Lin 等AAAI 2026
- Hidden in Plain Sight: Reasoning in Underspecified and Misspecified Scenarios for Multimodal LLMsQianqi Yan, Hongquan Li, Shan Jiang, Yang Zhao 等EMNLP 2025
- SkillGen: Learning Domain Skills for In-Context Sequential Decision MakingRuomeng Ding, Wei Cheng, Minglai Shao, Chen ZhaoAAAI 2026
