Lune

ACL2026顶会

HSCodeComp: A Realistic and Expert-level Agent Benchmark for Hierarchical Rule Application

Tian Lan, Yiqian Yang, Qianghuai Jia, Li Zhu, Hui Jiang, Hang Zhu, Weihua Luo, Longyue Wang

2026年份

摘要

Despite recent progress, existing agent benchmarks neglect a fundamental real-world capability: hierarchical rule application, a critical requirement in fields like law and medicine where agents must reason from broad categories down to specific exceptions. This introduces significant challenges, particularly in resolving complex logical dependencies and disambiguating fuzzy semantic boundaries. To bridge this gap, we introduce HSCODECOMP, an e-commerce benchmark that requires rigorous interpretation of product attributes and tariff rules to assign unique 10-digit Harmonized System Codes (HS Codes) to products 1 . HSCODECOMP comprises 632 products across 32 categories from e-commerce platforms. Each product contains detailed but noisy attributes mirroring real-world challenges, such as attributes and image. We also compile the official hierarchical tariff rules and knowledge databases. Utilizing these resources, 26 domain experts rigorously annotated the HSCodes. Evaluations of 23 state-of-the-art LLMs, VLMs, and agents reveal a large performance gap: the best-performing agent achieves only 46.8% accuracy versus 95.0% for human experts, and test-time scaling fails to close this gap. Extensive analysis reveals the challenges of hierarchical rule application. For example, excessive reasoning steps often induce "reasoning drift" that degrades accuracy. Furthermore, agents often misapply rules due to insufficient domain knowledge, as well as reasoning hallucinations that lack factual grounding 2 .

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper12

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖