Team Ai
Datasetpublic

xiapk7/ANT-A-Benchmark

πŸ“Š ANT-A Benchmark This directory contains the ANT-A Benchmark in HuggingFace-compatible Parquet format, ready for upload to the HuggingFace Hub. ANT-A is a benchmark of 7 real-world annotation tasks (10,500 samples, 5 annotators each), covering text and multimodal modalities. All 5 annotations are preserved per sample β€” no forced majority vote β€” supporting research on soft labels and annotator disagreement. πŸ—‚οΈ Dataset Structure Each task is provided as two… See the full description on the dataset page: https://huggingface.co/datasets/xiapk7/ANT-A-Benchmark.

sourceHugging Faceapache-2.0updated 12d agoView on Hugging Face
0likes176downloads
Dataset Card

πŸ“Š ANT-A Benchmark

[image]

This directory contains the ANT-A Benchmark in HuggingFace-compatible Parquet format, ready for upload to the HuggingFace Hub.

ANT-A is a benchmark of 7 real-world annotation tasks (10,500 samples, 5 annotators each), covering text and multimodal modalities. All 5 annotations are preserved per sample β€” no forced majority vote β€” supporting research on soft labels and annotator disagreement.

πŸ—‚οΈ Dataset Structure

Each task is provided as two Parquet files (_en and _zh) containing 1,500 samples each.

benchmark-huggingface/
β”œβ”€β”€ billing_scenario_classification/
β”‚   β”œβ”€β”€ billing_scenario_classification_en.parquet
β”‚   └── billing_scenario_classification_zh.parquet
β”œβ”€β”€ disease_privacy_assessment/
β”‚   β”œβ”€β”€ disease_privacy_assessment_en.parquet
β”‚   └── disease_privacy_assessment_zh.parquet
β”œβ”€β”€ entity_extraction/
β”‚   β”œβ”€β”€ entity_extraction_en.parquet
β”‚   └── entity_extraction_zh.parquet
β”œβ”€β”€ social_text_classification/
β”‚   β”œβ”€β”€ social_text_classification_en.parquet
β”‚   └── social_text_classification_zh.parquet
β”œβ”€β”€ image_classification/
β”‚   β”œβ”€β”€ image_classification_en.parquet
β”‚   β”œβ”€β”€ image_classification_zh.parquet
β”‚   └── images/                    # Image files referenced by image_url
β”‚       β”œβ”€β”€ d04483052ca827b9a8fb7f6d038284dc.jpg
β”‚       └── ...
β”œβ”€β”€ image_text_relevance/
β”‚   β”œβ”€β”€ image_text_relevance_en.parquet
β”‚   β”œβ”€β”€ image_text_relevance_zh.parquet
β”‚   └── images/                    # Image files referenced by image_url
β”‚       β”œβ”€β”€ f8bbec0e98904e16a740a519cf307ce1.jpg
β”‚       └── ...
└── image_attribution_recognition/
    β”œβ”€β”€ image_attribution_recognition_en.parquet
    β”œβ”€β”€ image_attribution_recognition_zh.parquet
    └── images/                    # Image files referenced by image_url
        β”œβ”€β”€ 03bf28fe75314bbe94e00e1dbc535446.jpg
        └── ...

Column Schema

ColumnTypeDescription
querystringThe input query (text prompt)
image_urlstring (vision tasks only)Relative path to the image file (e.g., images/abc123.jpg)
annotation_1 ~ annotation_5string5 independent annotator labels, JSON-encoded for structured tasks

Task Overview

TaskModalityDescriptionOutput TypeAgreement (5/5)
billing_scenario_classificationtextBilling scenario classificationsingle label90.1%
disease_privacy_assessmenttextDisease privacy compliance assessmentbinary (Yes/No)52.4%
entity_extractiontextMedical entity extractionlist of typed spans45.2%
social_text_classificationtextSocial media text classificationsingle label21.7%
image_classificationimageCategory / subcategory / guidance analysishierarchical multi-label35.5%
image_text_relevanceimage + textImage–text relevance assessmentordinal score75.9%
image_attribution_recognitionimageImage attribute recognitionbinary56.4%

Annotator Agreement Distribution

TaskFull agreement (5/5)Majority (3–4/5)No majority
billing_scenario_classification90.1%9.7%0.2%
image_text_relevance75.9%23.9%0.3%
image_attribution_recognition56.4%42.5%1.1%
disease_privacy_assessment52.4%47.5%0.1%
entity_extraction45.2%37.4%17.4%
image_classification35.5%50.3%14.2%
social_text_classification21.7%51.2%27.1%

πŸš€ Quick Start

python
from datasets import load_dataset

# Load from local parquet files
ds = load_dataset("parquet", data_files="billing_scenario_classification_en.parquet")["train"]

# Or load from HuggingFace Hub (replace with actual repo_id)
ds = load_dataset("xiapk7/ANT-A-Benchmark", "billing_scenario_classification")

For vision tasks, access images via image_url:

python
from PIL import Image
from pathlib import Path

ds = load_dataset("parquet", data_files="image_classification_en.parquet")["train"]
img_url = ds[0]["image_url"]  # "images/abc123.jpg"
img = Image.open(Path("image_classification") / img_url)

πŸ“ Data Format Details

Text Tasks

billing_scenario_classification

  • β€”query: Dict with keys Merchant Name, Invoice Title, Counterparty Name (JSON-encoded string)
  • β€”annotation_1~annotation_5: Single label string

disease_privacy_assessment

  • β€”query: Medical query text
  • β€”annotation_1~annotation_5: "Yes" or "No"

entity_extraction

  • β€”query: Medical text to extract entities from
  • β€”annotation_1~annotation_5: JSON list of {"text": "...", "category": "..."} objects

social_text_classification

  • β€”query: Social media post content
  • β€”annotation_1~annotation_5: Single label string

Vision / Multimodal Tasks

image_classification

  • β€”image_url: Relative path to the image
  • β€”annotation_1~annotation_5: JSON dict with Primary Category, Secondary Category, Discourse Analysis

image_text_relevance

  • β€”query: Text description
  • β€”image_url: Relative path to the image
  • β€”annotation_1~annotation_5: Ordinal score (e.g., "1 point", "2 points")

image_attribution_recognition

  • β€”image_url: Relative path to the image
  • β€”annotation_1~annotation_5: Binary string "0" or "1"

πŸ“‹ Annotation Rules

Each task includes its original annotation guideline (*_rules_en.md / *_rules_zh.md) shipped alongside the data. These are real production-grade rule documents (34–467 lines) that define the annotation schema and serve as prompts for rule-grounding research.


ℹ️ Notes

  • β€”Language: Original annotations are in Chinese; English translations are provided for the *_en files.
  • β€”Privacy: All personal-privacy data has been anonymized.
  • β€”Soft labels: No majority-vote aggregation is applied. All 5 annotator labels are preserved verbatim for research on label distribution and disagreement.
  • β€”Images: For the 3 vision tasks, images are stored in the images/ subdirectory and referenced by image_url in the parquet files.