Team Ai
Datasetpublic

CooperBench/qwen35-9b-git-coop

What this is Cooperative two-agent coding dataset: 211 task pairs across 18 repos, generated with mini_swe_agent on Qwen/Qwen3.5-9B in coop setting with a shared read-only git remote (--git). Agents coordinate via messaging and git fetch team; patches are auto-merged after both submit. At a glance Field Value Model Qwen/Qwen3.5-9B Agent mini_swe_agent (step_limit=300) Setting coop + git remote Repos 18 Pairs 211 Both-pass 5.7% (12/210… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen35-9b-git-coop.

sourceHugging Faceupdated 4mo agoView on Hugging Face
0likes136downloads
Dataset Card

What this is

Cooperative two-agent coding dataset: 211 task pairs across 18 repos, generated with mini_swe_agent on Qwen/Qwen3.5-9B in coop setting with a shared read-only git remote (--git). Agents coordinate via messaging and git fetch team; patches are auto-merged after both submit.

At a glance

FieldValue
ModelQwen/Qwen3.5-9B
Agentminisweagent (step_limit=300)
Settingcoop + git remote
Repos18
Pairs211
Both-pass5.7% (12/210 evaluated)
Per-feature pass16.0% (67/420)
Merge clean rate57.1% (120/210)
Total tokens~116.1M (in+out, from traj files)
OwnerArya Prabhudesai

How it was generated

bash
cooperbench run --setting coop --git -a mini_swe_agent -c 30 qwen35-9b-git-coop

Model served via vLLM OpenAI-compatible endpoint (openai/Qwen/Qwen3.5-9B).

File layout

  • —index.csv — one row per task pair; HF Dataset Viewer entry point
  • —qwen35-9b-git-coop/coop/<repo>/<task_id>/<features>/ — raw per-pair artifacts: result.json, eval.json, agent1_traj.json, agent2_traj.json, agent{1,2}.patch, conversation.json

log_dir column in index.csv points to the per-pair subdirectory.

Schema highlights for mid-training

Filter on: both_passed=true, model, agent_framework.

metadata JSON carries per-agent statuses, steps, merge outcome, per-feature pass — use json.loads(row["metadata"]) without following the pointer.

Note: total_tokens is 0 for this run — token counts are in agent_full_traj.json under messages[*].extra.response.usage (~116.1M total in+out).

Caveats

  • —step_limit=300 (3× default); 14.2% LimitsExceeded exits (60/422 agent slots) still present
  • —42.9% merge conflict rate (90/210) — conflicts skip evaluation
  • —Token fields in result.json are 0; aggregate from agent_full_traj.json if needed

Citation

bibtex
@dataset{qwen35_9b_git_coop,
  title  = {qwen35-9b-git-coop},
  author = {Arya Prabhudesai},
  year   = {2026},
  url    = {https://huggingface.co/datasets/CooperBench/qwen35-9b-git-coop}
}