HSCodeComp: A Realistic and Expert-level Agent Benchmark for Hierarchical Rule Application
Tian Lan, Yiqian Yang, Qianghuai Jia, Li Zhu, Hui Jiang, Hang Zhu, Weihua Luo, Longyue Wang
Abstract
Despite recent progress, existing agent benchmarks neglect a fundamental real-world capability: hierarchical rule application, a critical requirement in fields like law and medicine where agents must reason from broad categories down to specific exceptions. This introduces significant challenges, particularly in resolving complex logical dependencies and disambiguating fuzzy semantic boundaries. To bridge this gap, we introduce HSCODECOMP, an e-commerce benchmark that requires rigorous interpretation of product attributes and tariff rules to assign unique 10-digit Harmonized System Codes (HS Codes) to products 1 . HSCODECOMP comprises 632 products across 32 categories from e-commerce platforms. Each product contains detailed but noisy attributes mirroring real-world challenges, such as attributes and image. We also compile the official hierarchical tariff rules and knowledge databases. Utilizing these resources, 26 domain experts rigorously annotated the HSCodes. Evaluations of 23 state-of-the-art LLMs, VLMs, and agents reveal a large performance gap: the best-performing agent achieves only 46.8% accuracy versus 95.0% for human experts, and test-time scaling fails to close this gap. Extensive analysis reveals the challenges of hierarchical rule application. For example, excessive reasoning steps often induce "reasoning drift" that degrades accuracy. Furthermore, agents often misapply rules due to insufficient domain knowledge, as well as reasoning hallucinations that lack factual grounding 2 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on12
- WebShop: Towards Scalable Real-World Web Interaction with Grounded Language AgentsShunyu Yao, Howard Chen, John Yang, Karthik NarasimhanNeurIPS 2022 · 1,477 citations
- WebArena: A Realistic Web Environment for Building Autonomous AgentsShuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou et al.ICLR 2024 · 1,197 citations
- ALFWorld: Aligning Text and Embodied Environments for Interactive LearningMohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk et al.ICLR 2021 · 819 citations
- GAIA: a benchmark for General AI AssistantsGrégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun et al.ICLR 2024 · 716 citations
- OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task AutomationMengkang Hu, Yuhang Zhou, Wendong Fan, Yuzhou Nie et al.NeurIPS 2025 · 158 citations
Related papers
- HSGraphAgent: Knowledge-Graph-Guided Large Language Models for Harmonized System Code ClassificationQiang Xia, Zijian Zhang, Ao Wang, Wenhan Wang et al.ACL 2026
- FinSearchComp: Towards a Realistic, Expert-Level Evaluation of Financial Search and ReasoningLiang Hu, Jianpeng Jiao, Jiashuo Liu, Dongyuan Mutu et al.ICLR 2026 · 29 citations
- ACADREASON: Exploring the Limits of Reasoning Models with Academic Research ProblemsXin Gui, King Zhu, JinCheng Ren, Qianben Chen et al.ICLR 2026 · 1 citation
- DA-Code: Agent Data Science Code Generation Benchmark for Large Language ModelsYiming Huang, Jianwen Luo, Yan Yu, Yitong Zhang et al.EMNLP 2024 · 7 citations
- Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent EvaluationSayash Kapoor, Benedikt Stroebl, Peter Kirgis, Nitya Nadgir et al.ICLR 2026 · 86 citations
