HSCodeComp: A Realistic and Expert-level Agent Benchmark for Hierarchical Rule Application
Tian Lan, Yiqian Yang, Qianghuai Jia, Li Zhu, Hui Jiang, Hang Zhu, Weihua Luo, Longyue Wang
摘要
Despite recent progress, existing agent benchmarks neglect a fundamental real-world capability: hierarchical rule application, a critical requirement in fields like law and medicine where agents must reason from broad categories down to specific exceptions. This introduces significant challenges, particularly in resolving complex logical dependencies and disambiguating fuzzy semantic boundaries. To bridge this gap, we introduce HSCODECOMP, an e-commerce benchmark that requires rigorous interpretation of product attributes and tariff rules to assign unique 10-digit Harmonized System Codes (HS Codes) to products 1 . HSCODECOMP comprises 632 products across 32 categories from e-commerce platforms. Each product contains detailed but noisy attributes mirroring real-world challenges, such as attributes and image. We also compile the official hierarchical tariff rules and knowledge databases. Utilizing these resources, 26 domain experts rigorously annotated the HSCodes. Evaluations of 23 state-of-the-art LLMs, VLMs, and agents reveal a large performance gap: the best-performing agent achieves only 46.8% accuracy versus 95.0% for human experts, and test-time scaling fails to close this gap. Extensive analysis reveals the challenges of hierarchical rule application. For example, excessive reasoning steps often induce "reasoning drift" that degrades accuracy. Furthermore, agents often misapply rules due to insufficient domain knowledge, as well as reasoning hallucinations that lack factual grounding 2 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- WebShop: Towards Scalable Real-World Web Interaction with Grounded Language AgentsShunyu Yao, Howard Chen, John Yang, Karthik NarasimhanNeurIPS 2022 · 被引用 1,477 次
- WebArena: A Realistic Web Environment for Building Autonomous AgentsShuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou 等ICLR 2024 · 被引用 1,197 次
- ALFWorld: Aligning Text and Embodied Environments for Interactive LearningMohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk 等ICLR 2021 · 被引用 819 次
- GAIA: a benchmark for General AI AssistantsGrégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun 等ICLR 2024 · 被引用 716 次
- OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task AutomationMengkang Hu, Yuhang Zhou, Wendong Fan, Yuzhou Nie 等NeurIPS 2025 · 被引用 158 次
相关 Paper
- HSGraphAgent: Knowledge-Graph-Guided Large Language Models for Harmonized System Code ClassificationQiang Xia, Zijian Zhang, Ao Wang, Wenhan Wang 等ACL 2026
- FinSearchComp: Towards a Realistic, Expert-Level Evaluation of Financial Search and ReasoningLiang Hu, Jianpeng Jiao, Jiashuo Liu, Dongyuan Mutu 等ICLR 2026 · 被引用 29 次
- ACADREASON: Exploring the Limits of Reasoning Models with Academic Research ProblemsXin Gui, King Zhu, JinCheng Ren, Qianben Chen 等ICLR 2026 · 被引用 1 次
- DA-Code: Agent Data Science Code Generation Benchmark for Large Language ModelsYiming Huang, Jianwen Luo, Yan Yu, Yitong Zhang 等EMNLP 2024 · 被引用 7 次
- Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent EvaluationSayash Kapoor, Benedikt Stroebl, Peter Kirgis, Nitya Nadgir 等ICLR 2026 · 被引用 86 次
