Team Ai
Datasetpublic

PhilipZhang/AsynCodeBench

AsynCodeBench AsynCodeBench evaluates coding agents under five execution protocols while holding the task set and evaluation contracts fixed. The stable v0.4.2 release contains 19 coding tasks. Start with the tasks configuration; each row is one real benchmark task. Browse the 19 task examples in Dataset Viewer → The executable harness, manifests, container references, evaluation commands, and Quick Start are versioned in the official GitHub repository. Source repositories are… See the full description on the dataset page: https://huggingface.co/datasets/PhilipZhang/AsynCodeBench.

sourceHugging Faceotherupdated 6d agoView on Hugging Face
1likes85downloads
Dataset Card

AsynCodeBench

AsynCodeBench evaluates coding agents under five execution protocols while holding the task set and evaluation contracts fixed. The stable v0.4.2 release contains 19 coding tasks. Start with the tasks configuration; each row is one real benchmark task.

**Browse the 19 task examples in Dataset Viewer →**

The executable harness, manifests, container references, evaluation commands, and Quick Start are versioned in the official GitHub repository. Source repositories are materialized at each row's base_revision inside the recorded container image; the Parquet files therefore do not duplicate complete source checkouts or container layers.

Start here: one task example

The first tasks row is asyncodebench:cachetools. It contains the full coding prompt in problem_statement, the immutable source revision, evaluator command, test targets, and pinned container image/digest.

python
from datasets import load_dataset

tasks = load_dataset(
    "PhilipZhang/AsynCodeBench",
    "tasks",
    split="test",
    revision="v0.4.2",
)
example = tasks[0]
print(example["task_id"])             # asyncodebench:cachetools
print(example["problem_statement"])   # complete task prompt
print(example["container_image"])     # docker.io/wentingzhao/cachetools:v0
print(example["dependency_count"])    # 5

What the three configurations mean

ConfigurationRowsPurpose
`tasks`19Primary dataset. One row per coding task: prompt, source revision, tests, evaluator, and container. Most users only need this.
scenarios95Auxiliary protocol table: 19 tasks × 5 execution protocols. These are run conditions, not 95 additional coding tasks.
dependencies52Auxiliary annotation table: producer-consumer contracts inside the 19 tasks, with probe tests for ADPR, DRS, CAIL, and SAD metrics. These are not additional tasks.

All configurations use the test split because every row belongs to the evaluation benchmark.

Load auxiliary analysis tables

You do not need these tables to enumerate the 19 coding tasks. Load them when comparing protocols or computing dependency-aware metrics:

python
scenarios = load_dataset("PhilipZhang/AsynCodeBench", "scenarios", split="test", revision="v0.4.2")
dependencies = load_dataset("PhilipZhang/AsynCodeBench", "dependencies", split="test", revision="v0.4.2")

Run the benchmark

Follow the GitHub Quick Start for environment setup, model serving, official protocol commands, and result validation. Use the exact container image and digest recorded in the tasks configuration.

The five protocols are single, serial_specialists, async_private, caid_manager (displayed as Async-RO-Manager), and async_manager.

Version and scope

  • —Dataset release: v0.4.2
  • —GitHub revision: 2a639b049f249e79eb2031876bc06a110be9ac3a
  • —Release status: stable
  • —Official task count: 19
  • —Official protocol conditions: 95
  • —Official dependency points: 52

Candidate manifests outside the v0.4 release index are excluded. In particular, apache-tvm-20409 is not part of this dataset release.

License

The original AsynCodeBench material is Apache-2.0. Referenced upstream tasks and repositories retain their own licenses. See LICENSE and THIRD_PARTY_NOTICES.md.

Citation

If you use AsynCodeBench, cite the software release using the repository's CITATION.cff.