PhilipZhang/AsynCodeBench
AsynCodeBench AsynCodeBench evaluates coding agents under five execution protocols while holding the task set and evaluation contracts fixed. The stable v0.4.2 release contains 19 coding tasks. Start with the tasks configuration; each row is one real benchmark task. Browse the 19 task examples in Dataset Viewer → The executable harness, manifests, container references, evaluation commands, and Quick Start are versioned in the official GitHub repository. Source repositories are… See the full description on the dataset page: https://huggingface.co/datasets/PhilipZhang/AsynCodeBench.
AsynCodeBench
AsynCodeBench evaluates coding agents under five execution protocols while holding the task set and evaluation contracts fixed. The stable v0.4.2 release contains 19 coding tasks. Start with the tasks configuration; each row is one real benchmark task.
**Browse the 19 task examples in Dataset Viewer →**
The executable harness, manifests, container references, evaluation commands, and Quick Start are versioned in the official GitHub repository. Source repositories are materialized at each row's base_revision inside the recorded container image; the Parquet files therefore do not duplicate complete source checkouts or container layers.
Start here: one task example
The first tasks row is asyncodebench:cachetools. It contains the full coding prompt in problem_statement, the immutable source revision, evaluator command, test targets, and pinned container image/digest.
from datasets import load_dataset
tasks = load_dataset(
"PhilipZhang/AsynCodeBench",
"tasks",
split="test",
revision="v0.4.2",
)
example = tasks[0]
print(example["task_id"]) # asyncodebench:cachetools
print(example["problem_statement"]) # complete task prompt
print(example["container_image"]) # docker.io/wentingzhao/cachetools:v0
print(example["dependency_count"]) # 5What the three configurations mean
All configurations use the test split because every row belongs to the evaluation benchmark.
Load auxiliary analysis tables
You do not need these tables to enumerate the 19 coding tasks. Load them when comparing protocols or computing dependency-aware metrics:
scenarios = load_dataset("PhilipZhang/AsynCodeBench", "scenarios", split="test", revision="v0.4.2")
dependencies = load_dataset("PhilipZhang/AsynCodeBench", "dependencies", split="test", revision="v0.4.2")Run the benchmark
Follow the GitHub Quick Start for environment setup, model serving, official protocol commands, and result validation. Use the exact container image and digest recorded in the tasks configuration.
The five protocols are single, serial_specialists, async_private, caid_manager (displayed as Async-RO-Manager), and async_manager.
Version and scope
- Dataset release:
v0.4.2 - GitHub revision:
2a639b049f249e79eb2031876bc06a110be9ac3a - Release status: stable
- Official task count: 19
- Official protocol conditions: 95
- Official dependency points: 52
Candidate manifests outside the v0.4 release index are excluded. In particular, apache-tvm-20409 is not part of this dataset release.
License
The original AsynCodeBench material is Apache-2.0. Referenced upstream tasks and repositories retain their own licenses. See LICENSE and THIRD_PARTY_NOTICES.md.
Citation
If you use AsynCodeBench, cite the software release using the repository's CITATION.cff.
