ATH-MaaS/HSCodeComp
HSCodeComp: A Realistic and Expert-Level Benchmark for Deep Search Agents in Hierarchical Rule Application Paper | Code | Dataset on Hugging Face ⭐ MarcoPolo Team ⭐ Alibaba Group 🗂️ Data 📌 Overview HSCodeComp is the first realistic, expert-level e-commerce benchmark designed to evaluate deep search agents on their ability to perform Level-3 knowledge—hierarchical rule application—a critical yet overlooked capability in current agent evaluation… See the full description on the dataset page: https://huggingface.co/datasets/ATH-MaaS/HSCodeComp.
12625
1---2license: apache-2.03task_categories:4 - text-classification5 - question-answering6tags:7 - e-commerce8 - agentic-ai9 - code-classification10 - rule-application11 - benchmark12---13 14# HSCodeComp: A Realistic and Expert-Level Benchmark for Deep Search Agents in Hierarchical Rule Application15 16[Paper](https://arxiv.org/abs/2510.19631) | [Code](https://github.com/AIDC-AI/Marco-Search-Agent/tree/main/HSCodeComp) | [Dataset on Hugging Face](https://huggingface.co/datasets/AIDC-AI/HSCodeComp)17 18[](https://opensource.org/licenses/Apache-2.0)19[](https://www.python.org/downloads/)20[](https://huggingface.co/datasets/AIDC-AI/HSCodeComp)21 22<div align="center">23 24⭐ _**MarcoPolo Team**_ ⭐25 26[_**Alibaba Group**_](https://qianwenai.com)27 28🗂️ [**Data**](https://github.com/AIDC-AI/Marco-Search-Agent/tree/main/HSCodeComp/data/test_data.jsonl)29 30</div>31 32---33 34## 📌 Overview35 36 37<div align="center">38 <img src="assets/overview.png" alt="Overview" width="60%" style="display: inline-block; vertical-align: top; margin-right: 2%;">39 <img src="assets/teaser_img.png" alt="Teaser" width="34%" style="display: inline-block; vertical-align: top;">40</div>41 42**HSCodeComp** is the first realistic, expert-level e-commerce benchmark designed to evaluate **deep search agents** on their ability to perform Level-3 knowledge—**hierarchical rule application**—a critical yet overlooked capability in current agent evaluation frameworks.43 44The task requires agents to predict the exact **10-digit Harmonized System Code (HSCode)** for products described with **noisy, real-world e-commerce domain**, by correctly applying complex, hierarcahical tariff rules (e.g., from eWTP and official customs rulings). These rules often contain **vague language** and **implicit logic**, making accurate classification highly challenging. Our evaluation reveals a stark performance gap:45 46* 🔹 **Best AI agent (SmolAgent + GPT-5 VLM): 46.8%**47* 🔹 **Human experts: 95.0%**48 49Besides, ablation study also reveals that **inference-time scaling fails to improve the performance**. These highlight that deep search with **hierarchical rule application** remains a major unsolved challenge for state-of-the-art AI agent systems. 50 51---52 53## 🔥 News54* [2025/10/] 🔥 We released the [paper](https://arxiv.org/abs/2510.19631) and [dataset](https://huggingface.co/datasets/AIDC-AI/HSCodeComp) of our challenging HSCodeComp dataset.55 56---57 58## 📋 Dataset59 6061This figure reveals that the data format HSCodeComp dataset.62 63### Input64 65Each product $x \in \mathcal{X}$ contains rich information: $x = (t, A, c, i, p, u, r)$, where:66- **$t$**: Product title67- **$A = \{(k_j, v_j)\}_{j=1}^K$**: Set of $K$ product attributes (e.g., material, package size)68- **$c$**: Product categories defined by the e-commerce platform69- **$p$**: Price70- **$u$**: Currency71 72### Knowledge: Hierarchical Rules73The task requires agents to effectively utilize three types of e-commerce domain knowledge:741. **Hierarchical tariff rules** from official classification systems (e.g., eWTP) with complex implicit logic and vague linguistic constraints752. **Human-written decision rules** that specify how to correctly apply tariff rules763. **Official customs rulings databases** (e.g., U.S. CROSS) containing historical HSCode classification decisions77 78### Output79 80The HSCode $y \in \mathcal{Y}$ is a single **10-digit numeric string** $\mathcal{Y} \subseteq \{0,1,\ldots,9\}^{10}$. The HSCode structure is hierarchical:81- **First 2 digits**: HS chapter82- **First 4 digits**: HS heading 83- **First 6 digits**: HS sub-heading84- **Last 4 digits (7-10)**: Country-specific codes85 86The 10-digit HSCode must follow a valid path in the official HS taxonomy. Please refer to [our paper](https://arxiv.org/abs/2510.19631) for more details about these data.87 88### Dataset Collection and Statistic89 9091We engage several domain experts in HSCode prediction, and conduct a well-designed 6 steps pipeline to construct dataset. The important details of our proposed HSCodeComp is provided in following table.92 93| Metric | Value |94|--------|-------|95| **Total Products** | 632 expert-annotated entries |96| **HS Chapters** | 27 chapters |97| **First-level Categories** | 32 categories |98| **Data Source** | Large-scale e-commerce platforms |99| **Validation** | Multiple domain experts |100| **Inter-annotator Agreement** | >98% |101| **Models Tested** | 14 foundation models, 6 open-source agents, 3 closed-source systems |102| **Knowledge Level** | Level 3: Hierarchical rule application |103 104 105---106 107## ⚙️ Sample Usage108 109### 📁 Repository Structure110 111```bash112HSCodeComp/113├── data/114│ └── test_data.csv # Product descriptions, attributes and ground-truth HSCodes115├── eval/116│ └── test_llm.py # Evaluation script for model predictions117├── LICENSE118└── README.md119```120 121### 🛠️ Environment Setup122 123```bash124 125# Create and activate a virtual environment (optional but recommended)126python -m venv hscodcomp_env127source hscodcomp_env/bin/activate # Linux/macOS128# hscodcomp_env\Scripts\activate # Windows129 130# Install dependencies (e.g., pandas, etc.)131pip install pandas,openai,tqdm,threading,dotenv132 133# set openai keys and base urls in HSCodeComp/.env134```135 136### 🚀 Run Evaluation137 138```bash139# Set models_to_test = ["gpt-4o"] in eval/test_llm.py140python eval/test_llm.py141```142 143The script reports **exact-match accuracy** at **2-digit, 4-digit, 6-digit, 8-digit, and 10-digit** levels.144 145---146 147## 📊 Benchmark Performance148 149### Complete Evaluation on HSCodeComp150151 152The top-performming baseline SmolAgent (GPT-5 with vision capability) achieves the best performance, while it sill largely lag behind human expert performance.153 154<img src="assets/closed_source_main_exp_result.jpg" alt="Closed Source Main Exp Result" width="60%" style="border-radius: 8px; box-shadow: 0 2px 8px rgba(0,0,0,0.1);">155 156Closed-source agent systems still largely underperform domain expert and open-source agent systems with GPT-5 backbone model.157 158### Current Agents Fail to Leverage Hierarchical Decision Rules159 160<img src="assets/fail_to_use_DR.jpg" alt="Closed Source Main Exp Result" width="75%" style="border-radius: 8px; box-shadow: 0 2px 8px rgba(0,0,0,0.1);">161 162Performance degrades when human decision rules are included in the system prompt.163 164 165### More Thinking Leads to Worse Performance166<img src="assets/more_think_leads_to_worse_performance.jpg" alt="Closed Source Main Exp Result" width="75%" style="border-radius: 8px; box-shadow: 0 2px 8px rgba(0,0,0,0.1);">167 168* More thinking leads to more errors and hallucinations in this highly domain-specific HSCode prediction task.169* When accurate information is available, through calling tools, prioritizing tool utilization over reasoning yields better results.170 171### Test-time Scaling Fails to Improve Performance172<img src="assets/tts.jpg" alt="Closed Source Main Exp Result" width="80%" style="border-radius: 8px; box-shadow: 0 2px 8px rgba(0,0,0,0.1);">173 174Two kinds of inference-time scaling strategy (majority voting and self-reflection) fails to effectively improve the performance.175 176> For complete experimental results, please refer to [our paper](https://arxiv.org/abs/2510.19631).177 178---179 180## 🤝 Acknowledgements181 182We thank the human experts who meticulously annotated and validated the HSCodes. Their domain knowledge is the foundation of this benchmark’s quality and realism.183 184---185 186## 🛡️ License187 188This project is licensed under the **Apache-2.0 License**189 190---191 192## ⚠️ DISCLAIMER193Our datasets are constructed using publicly accessible product data sources. Although we remove the product image and url in the HSCodeComp, we still cannot guarantee that our datasets are completely free of copyright issues or improper content. If you believe anything infringes on your rights or generates improper content, please contact us ([Tian Lan](https://github.com/gmftbyGMFTBY) and [Longyue Wang](https://www.longyuewang.com/)), and we will promptly address the matter.194 195---