Team Ai
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01OpenMOSS-Team /SWE-bench-Science SWE-bench Science SWE-bench Science evaluates coding agents on software-engineering tasks drawn from scientific-computing repositories. The release contains 119 tasks across 20 scientific domains, with isolated environments and separate programmatic verifiers. GitHub release repository: OpenMOSS/SWE-bench-Science Runtime images: Docker Hub, pinned by immutable linux/amd64 digests Evaluation framework: Pier, compatible with Harbor task format Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/SWE-bench-Science.textn<1K9 likes2.1k downloads9d agoHugging Face02UnipatAI /Monthly-SWEBench-2026-05 Monthly-SWEBench 2026-05 This package contains the 2026-05 Monthly-SWEBench final release set. It includes 100 Harbor-format software engineering tasks selected from closed GitHub PRs and validated with oracle=1 / nop=0. Files bugfix.tar.zst: 50 bug-oriented repair or maintenance tasks. non_bugfix.tar.zst: 50 feature, API evolution, or engineering-improvement tasks. preview.csv: task ids, split labels, source change buckets, and archive paths. tasks.conf: one… See the full description on the dataset page: https://huggingface.co/datasets/UnipatAI/Monthly-SWEBench-2026-05.texttext-generationn<1K0 likes1.6k downloads4mo agoHugging Face03UnipatAI /Monthly-SWEBench-2026-04 Monthly-SWEBench-2026-04 Monthly-SWEBench-2026-04 is a curated benchmark of 90 real-world software engineering tasks, sourced from GitHub pull requests merged in April 2026. Tasks are in Harbor format and can be run with any Harbor-compatible agent. View leaderboard and results → 90 tasks — 43 bugfix + 47 non-bugfix Tasks span diverse open-source repositories Each task includes a runnable environment, test suite, and reference solution Task Structure Each task is a… See the full description on the dataset page: https://huggingface.co/datasets/UnipatAI/Monthly-SWEBench-2026-04.texttext-generationn<1K0 likes1.5k downloads5mo agoHugging Face04UnipatAI /Monthly-SWEBench-2026-03 Monthly-SWEBench-2026-03 Monthly-SWEBench-2026-03 is a curated benchmark of 112 real-world software engineering tasks, sourced from GitHub pull requests merged in March 2026. Tasks are in Harbor format and can be run with any Harbor-compatible agent. View leaderboard and results → 112 tasks — 68 bugfix + 44 non-bugfix Tasks span diverse open-source repositories Each task includes a runnable environment, test suite, and reference solution Task Structure Each… See the full description on the dataset page: https://huggingface.co/datasets/UnipatAI/Monthly-SWEBench-2026-03.texttext-generationn<1K1 likes1.4k downloads4mo agoHugging Face05Contextbench /SWE-bench_Pro Dataset Summary SWE-Bench Pro is a challenging, enterprise-level dataset for testing agent ability on long-horizon software engineering tasks. Paper: https://static.scale.com/uploads/654197dc94d34f66c0f5184e/SWEAP_Eval_Scale%20(9).pdf See the related evaluation Github: https://github.com/scaleapi/SWE-bench_Pro-os Dataset Structure We follow SWE-Bench Verified (https://huggingface.co/datasets/SWE-bench/SWE-bench_Verified) in terms of dataset structure, with several… See the full description on the dataset page: https://huggingface.co/datasets/Contextbench/SWE-bench_Pro.textn<1K0 likes447 downloads10mo agoHugging Face06microsoft /SWE-Sharp-Bench SWE-Sharp-Bench SWE-Sharp-Bench is a comprehensive benchmark suite for evaluating software engineering capabilities of AI agents and models on C# and .NET codebases. This benchmark extends the SWE-Bench framework to the C# ecosystem, providing real-world software engineering tasks from popular open-source repositories. Code - https://github.com/microsoft/prose/tree/main/misc/SWE-Sharp-Bench Research Paper Draft & Benchmark Analysis: https://aka.ms/swesharparxiv Contact… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/SWE-Sharp-Bench.textn<1K3 likes411 downloads11mo agoHugging Face07swebenchatlas /swe-bench-atlas-anon-public SWE-bench Atlas 1. Summary In the domain of software engineering, LLM capabilities have progressed rapidly, underscoring the need for evolving evaluation frameworks. While foundational, benchmarks like SWE-bench, SWE-bench Verified, and other such variants are incomplete, with manually curated design causing scalability bottlenecks, weak test oracles, dataset aging and contamination, reproducibility challenges, and more. In response, we introduce SWE-bench Atlas: a… See the full description on the dataset page: https://huggingface.co/datasets/swebenchatlas/swe-bench-atlas-anon-public.textn<1K0 likes389 downloads11mo agoHugging Face08CharlieLLL /SWEbench-Verified-eval150-u355-M2.7-orch-cost-20260923 Fixed Solo350 u355 + MiniMax-M2.7: orchestration cost study Best observed cost tradeoff: compact coordinator decisions plus soft review at the existing hard limit (at most 12 worker turns). M2.7 metered token cost falls 59.3%, while mean solved tasks decrease from 90.00 to 87.67/150. Accuracy equivalence was not established. This closed study contains 6 designs and 16 complete independent runs on the same 150 tasks (2400 scored task/run pairs), each with an independent audit.… See the full description on the dataset page: https://huggingface.co/datasets/CharlieLLL/SWEbench-Verified-eval150-u355-M2.7-orch-cost-20260923.tabular1K<n<10K0 likes321 downloads17d agoHugging Face09CharlieLLL /SWEbench-Verified-eval150-M2.7-Qwen3.5-9B-orch-7arms-2repeats-w32-20260920 SWE-bench Verified eval150 — M2.7 × Qwen3.5-9B, seven arms, two repeats, 32 concurrency Campaign 2026-09-20. 14/14 independent full150 runs audited. Complete accuracy evidence. Evaluation mode is orch: MiniMax-M2.7 orchestrator and the specified Qwen3.5-9B worker. Training mode is labeled independently. All runs use 32 concurrent episodes, 10GiB Docker sandboxes, four TP1 workers and one TP4/EP4 coordinator. Frozen regression-gated prompts, decoding and canonical verifier match… See the full description on the dataset page: https://huggingface.co/datasets/CharlieLLL/SWEbench-Verified-eval150-M2.7-Qwen3.5-9B-orch-7arms-2repeats-w32-20260920.tabulartext-generation1K<n<10K0 likes154 downloads20d agoHugging Face10huyouare /SWE-bench_Verified_With_Annotationstabularn<1K1 likes94 downloads2y agoHugging Face11ldowey /swe_mon_benchmarktextn<1K0 likes89 downloads3y agoHugging Face12huyouare /SWE-bench_Verified_Smalltabularn<1K1 likes78 downloads2y agoHugging Face13hrw /SWE-bench_Litetextn<1K0 likes66 downloads2y agoHugging Face14bvisser /SWE-bump-benchtabularn<1K0 likes25 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.