xiapk7/ANT-A-Benchmark
π ANT-A Benchmark This directory contains the ANT-A Benchmark in HuggingFace-compatible Parquet format, ready for upload to the HuggingFace Hub. ANT-A is a benchmark of 7 real-world annotation tasks (10,500 samples, 5 annotators each), covering text and multimodal modalities. All 5 annotations are preserved per sample β no forced majority vote β supporting research on soft labels and annotator disagreement. ποΈ Dataset Structure Each task is provided as twoβ¦ See the full description on the dataset page: https://huggingface.co/datasets/xiapk7/ANT-A-Benchmark.
π ANT-A Benchmark
This directory contains the ANT-A Benchmark in HuggingFace-compatible Parquet format, ready for upload to the HuggingFace Hub.
ANT-A is a benchmark of 7 real-world annotation tasks (10,500 samples, 5 annotators each), covering text and multimodal modalities. All 5 annotations are preserved per sample β no forced majority vote β supporting research on soft labels and annotator disagreement.
ποΈ Dataset Structure
Each task is provided as two Parquet files (_en and _zh) containing 1,500 samples each.
benchmark-huggingface/
βββ billing_scenario_classification/
β βββ billing_scenario_classification_en.parquet
β βββ billing_scenario_classification_zh.parquet
βββ disease_privacy_assessment/
β βββ disease_privacy_assessment_en.parquet
β βββ disease_privacy_assessment_zh.parquet
βββ entity_extraction/
β βββ entity_extraction_en.parquet
β βββ entity_extraction_zh.parquet
βββ social_text_classification/
β βββ social_text_classification_en.parquet
β βββ social_text_classification_zh.parquet
βββ image_classification/
β βββ image_classification_en.parquet
β βββ image_classification_zh.parquet
β βββ images/ # Image files referenced by image_url
β βββ d04483052ca827b9a8fb7f6d038284dc.jpg
β βββ ...
βββ image_text_relevance/
β βββ image_text_relevance_en.parquet
β βββ image_text_relevance_zh.parquet
β βββ images/ # Image files referenced by image_url
β βββ f8bbec0e98904e16a740a519cf307ce1.jpg
β βββ ...
βββ image_attribution_recognition/
βββ image_attribution_recognition_en.parquet
βββ image_attribution_recognition_zh.parquet
βββ images/ # Image files referenced by image_url
βββ 03bf28fe75314bbe94e00e1dbc535446.jpg
βββ ...Column Schema
Task Overview
Annotator Agreement Distribution
π Quick Start
from datasets import load_dataset
# Load from local parquet files
ds = load_dataset("parquet", data_files="billing_scenario_classification_en.parquet")["train"]
# Or load from HuggingFace Hub (replace with actual repo_id)
ds = load_dataset("xiapk7/ANT-A-Benchmark", "billing_scenario_classification")For vision tasks, access images via image_url:
from PIL import Image
from pathlib import Path
ds = load_dataset("parquet", data_files="image_classification_en.parquet")["train"]
img_url = ds[0]["image_url"] # "images/abc123.jpg"
img = Image.open(Path("image_classification") / img_url)π Data Format Details
Text Tasks
billing_scenario_classification
query: Dict with keysMerchant Name,Invoice Title,Counterparty Name(JSON-encoded string)annotation_1~annotation_5: Single label string
disease_privacy_assessment
query: Medical query textannotation_1~annotation_5:"Yes"or"No"
entity_extraction
query: Medical text to extract entities fromannotation_1~annotation_5: JSON list of{"text": "...", "category": "..."}objects
social_text_classification
query: Social media post contentannotation_1~annotation_5: Single label string
Vision / Multimodal Tasks
image_classification
image_url: Relative path to the imageannotation_1~annotation_5: JSON dict withPrimary Category,Secondary Category,Discourse Analysis
image_text_relevance
query: Text descriptionimage_url: Relative path to the imageannotation_1~annotation_5: Ordinal score (e.g.,"1 point","2 points")
image_attribution_recognition
image_url: Relative path to the imageannotation_1~annotation_5: Binary string"0"or"1"
π Annotation Rules
Each task includes its original annotation guideline (*_rules_en.md / *_rules_zh.md) shipped alongside the data. These are real production-grade rule documents (34β467 lines) that define the annotation schema and serve as prompts for rule-grounding research.
βΉοΈ Notes
- Language: Original annotations are in Chinese; English translations are provided for the
*_enfiles. - Privacy: All personal-privacy data has been anonymized.
- Soft labels: No majority-vote aggregation is applied. All 5 annotator labels are preserved verbatim for research on label distribution and disagreement.
- Images: For the 3 vision tasks, images are stored in the
images/subdirectory and referenced byimage_urlin the parquet files.
