Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01harborframework /terminal-bench-2.1 Terminal-Bench 2.1 (Harbor git-repos dataset) Harbor website · Harbor GitHub This is a private mirror of the task content from harbor-framework/terminal-bench-2-1 at commit 7131e43 (the source repo has no tagged releases yet), laid out so it can be consumed directly by Harbor's git-repos dataset support. The primary source is the GitHub repository above — please open issues and pull requests there, not here. How to run Always pass the full URL, not org/name — a… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/terminal-bench-2.1.documentn<1K15 likes188k downloads22d agoHugging Face02harborframework /terminal-bench-3.0 Terminal-Bench 3.0 The primary source is hosted on GitHub, please open issues and pull requests there, not here. The official published dataset is hosted on the Harbor Hub along with the official leaderboard. Usage e.g. harbor run -d terminal-bench/terminal-bench@3.0.0 This repo is a mirror of harbor-framework/terminal-bench at tag v3.0.0, laid out so it can be consumed directly by Harbor's git-repos dataset support. How to run via this Huggingface repo Always… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/terminal-bench-3.0.6 likes169k downloads2mo agoHugging Face03harborframework /terminal-bench Terminal-Bench The primary source is hosted on GitHub, please open issues and pull requests there, not here. Terminal-Bench is now a continuous benchmark: new versions are released periodically as tags on the source repo instead of one-off snapshots. This dataset mirrors that model on the Hub: instead of a separate terminal-bench-X.Y repo per release, one repo, tagged per version. main always tracks the latest published version; each release is additionally available as an… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/terminal-bench.documentn<1K6 likes130k downloads22d agoHugging Face04harborframework /terminal-bench-2.0Warning: The leaderboard above is unofficial. The official leaderboard is https://www.tbench.ai/leaderboard/terminal-bench/2.0, in which entires are audited for correct configuration, results show which agent harness is used, and verified trajectories are publicly viewable. Warning: The dataset is a read-only mirror. The primary source for this dataset is on GitHub: https://github.com/harbor-framework/terminal-bench-2. Please open issues and pull requests there. How this mirror was created… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/terminal-bench-2.0.text-generationn<1K51 likes85k downloads5mo agoHugging Face05harborframework /terminal-bench-science Terminal-Bench-Science The primary source is hosted on GitHub, please open issues and pull requests there, not here. Terminal-Bench-Science is a benchmark of real-world computational research workflows across the life, physical, earth, mathematical, and engineering sciences. Like Terminal-Bench, it's a continuous benchmark: releases are published as tags on the source repo. This dataset mirrors that on the Hub: one repo, tagged per version, instead of a separate repo per… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/terminal-bench-science.textn<1K6 likes84k downloads22d agoHugging Face06harborframework /terminal-bench-science-lfs Terminal-Bench-Science — task input mirror Large input files for Terminal-Bench-Science tasks, which cannot be committed to git. Tasks pull from here at container build time, pinned to a commit SHA and verified against a checksum file that ships in the task directory. One top-level prefix per task; everything lives under <task-name>/input/. Benchmark contamination canary This dataset is benchmark material. If you are assembling a training corpus, exclude it.… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/terminal-bench-science-lfs.0 likes70k downloads1mo agoHugging Face07penfever /terminal-bench-2 Terminal-Bench-2.0 Beta Welcome to Terminal-Bench-2.0! If you’re reading this you’re a member of the Terminal-Bench community that we’ve selected to get a sneak peek at the latest version of the benchmark. Getting Started First, clone Harbor (formerly “Sandboxes”): git clone https://github.com/laude-institute/harbor.git From inside the Harbor directory run: uv sync This will install Harbor, our new package for running agent evals. You should now be able to run TB 2.0!… See the full description on the dataset page: https://huggingface.co/datasets/penfever/terminal-bench-2.documentn<1K1 likes24k downloads11mo agoHugging Face08zai-org /terminal-bench-2-verified Terminal-Bench 2.0 Verified: Instruction & Environment Fix Version 中文版本 2026.08.18 Update: Following the release of Terminal-Bench 2.1, we conducted another round of review on several tasks' grading and instructions, validating each fix end-to-end inside the actual task images. This round fixes 6 tasks in two categories: Grading/Test Fixes (dna-insert, make-doom-for-mips, filter-js-from-html, install-windows-3.11) — the test logic itself misjudged valid submissions or let… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/terminal-bench-2-verified.83 likes16k downloads1mo agoHugging Face09IntelligenceLab /Long-Horizon-Terminal-Bench Long-Horizon Terminal-Bench (LHTB) LHTB is a 46-task benchmark for measuring how well LLM agents sustain useful work in a containerized terminal over hundreds of steps. Unlike short-horizon coding benchmarks where an agent writes one artifact and stops, LHTB drops the agent into a stateful environment and grades it with hidden, rebuild-from-artifact verifiers — self-reported progress does not count. 📝 Blog: https://zli12321.github.io/LHTB/ 🏆 Leaderboard:… See the full description on the dataset page: https://huggingface.co/datasets/IntelligenceLab/Long-Horizon-Terminal-Bench.documenttext-generationn<1K137 likes12k downloads20d agoHugging Face10harborframework /terminal-bench-2-leaderboard Terminal-Bench 2.0 Leaderboard Submissions This repository accepts leaderboard submissions for Terminal-Bench 2.0. How to Submit Fork this repository Create a new branch for your submission Add your submission (a job or folder of jobs) under submissions/terminal-bench/2.0/<agent>__<model(s)>/ Open a Pull Request Submission Structure submissions/ terminal-bench/ 2.0/ <agent>__<model>/ metadata.yaml # Required: agent and model info… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/terminal-bench-2-leaderboard.36 likes9.2k downloads5mo agoHugging Face11DCAgent2 /terminal_bench_2 Terminal-Bench 2.0 ###################################################################### # _____ _ _ ______________ # # |_ _|__ _ __ _ __ ___ (_)_ __ __ _| | || || # # | |/ _ \ '__| '_ ` _ \| | '_ \ / _` | | || > || # # | | __/ | | | | | | | | | | | (_| | | || || # # |_|\___|_| |_| |_| |_|_|_| |_|\__,_|_| ||____________|| # # ____ _ ____… See the full description on the dataset page: https://huggingface.co/datasets/DCAgent2/terminal_bench_2.0 likes8.5k downloads11mo agoHugging Face12harborframework /terminal-bench-lfs0 likes6.9k downloads1mo agoHugging Face13alibabagroup /terminal-bench-pro Terminal-Bench Pro Overview Terminal-Bench Pro is a systematic extension of the original Terminal-Bench, designed to address key limitations in existing terminal-agent benchmarks. 400 tasks (200 public + 200 private) across 8 domains: data processing, games, debugging, system admin, scientific computing, software engineering, ML, and security Expert-designed tasks derived from real-world scenarios and GitHub issues High test coverage with ~28.3 test cases per… See the full description on the dataset page: https://huggingface.co/datasets/alibabagroup/terminal-bench-pro.texttext-generationn<1K5 likes5.9k downloads9mo agoHugging Face14ia03 /terminal-bench Terminal-Bench Dataset This dataset contains tasks from Terminal-Bench, a benchmark for evaluating AI agents in real terminal environments. Each task is packaged as a complete, self-contained archive that preserves the exact directory structure, binary files, Docker configurations, and test scripts needed for faithful reproduction. The archive column contains a gzipped tarball of the entire task directory. Dataset Overview Terminal-Bench evaluates AI agents on… See the full description on the dataset page: https://huggingface.co/datasets/ia03/terminal-bench.tabulartext-generationn<1K3 likes4k downloads1y agoHugging Face15yoonholee /terminalbench-trajectories Terminal-Bench 2.0 Trajectories Full agent trajectories from Terminal-Bench 2.0, a benchmark that evaluates AI coding agents on real-world terminal tasks. Each row is one trial: an agent attempting a task, with the complete step-by-step trace of messages, tool calls, and observations. Explorer: yoonholee.com/web-apps/terminal-bench Quick start from datasets import load_dataset import json ds = load_dataset("yoonholee/terminalbench-trajectories", split="train") #… See the full description on the dataset page: https://huggingface.co/datasets/yoonholee/terminalbench-trajectories.tabulartext-generation10K<n<100K17 likes2.8k downloads7mo agoHugging Face16Zhongzhi1228 /Terminal-Bench-Hard Terminal-Bench Hard Terminal-Bench Hard is a set of 100 challenging terminal-based agent tasks. The tasks cover software engineering, debugging, data processing, system administration, security, scientific computing, and related command-line workflows. Contents tasks/: runnable tasks in Harbor format. metadata/tasks.parquet: searchable task metadata and instructions. Each task directory contains task.toml, instruction.md, an environment/ directory, and verifier… See the full description on the dataset page: https://huggingface.co/datasets/Zhongzhi1228/Terminal-Bench-Hard.imagequestion-answeringn<1K0 likes2.6k downloads2mo agoHugging Face17LocalLLaMA /terminal-bench-mini terminal-bench-mini Fourteen of Terminal-Bench 2.0's ninety tasks, picked so that ranking agents on the subset reproduces ranking them on the whole benchmark. Running ninety tasks five times each is how the official leaderboard is built. That is out of reach if you are comparing quant variants, fine-tunes or local models on your own hardware. This subset turns a multi-day sweep into a few hours. Same approach as deepswe-mini: take the published per-task results, rank the field… See the full description on the dataset page: https://huggingface.co/datasets/LocalLLaMA/terminal-bench-mini.tabularn<1K6 likes2.2k downloads16d agoHugging Face18mlfoundations-dev /terminal-bench-traces-localtext1K<n<10K0 likes2.1k downloads1y agoHugging Face19Lottolabs /terminal-bench-2.1-qwen3.8-27b-traces Terminal-Bench 2.1 traces: Qwen3.8-27B-GPTQ-4bit, xhigh / medium / low / off Complete agent trajectories, verifier output, timing and token usage for all 89 Terminal-Bench 2.1 tasks run locally with btbtyler09/Qwen3.8-27B-GPTQ-4bit on 2× RTX 3090, plus the adaptive fallback reruns at lower reasoning effort. Headline result: 62/89 (69.66%) at xhigh in a single clean pass. Cumulative best-of across xhigh → medium → low → off fallbacks: 70/89 (78.65%). The second number is not a… See the full description on the dataset page: https://huggingface.co/datasets/Lottolabs/terminal-bench-2.1-qwen3.8-27b-traces.text1K<n<10K1 likes2k downloads1mo agoHugging Face20pimpalgaonkar /terminalbench-sqlite-db0 likes1.8k downloads1y agoHugging Face21laion /terminal-bench-2-1-apptainer-v1 Terminal-Bench 2.1 (offline Apptainer, v1) A validated subset of Terminal-Bench 2.1 (revision 7131e4375048a0e408a8fb404b5f499d726b695b, Apache-2.0) for running on HPC clusters without Docker and without internet on compute nodes, with the harbor Apptainer bridge. 68 of the included tasks are byte-identical to upstream; 6 carry local changes (see below). Scores on this subset are not comparable to full 89-task TB2.1 leaderboard numbers. status tasks meaning validated… See the full description on the dataset page: https://huggingface.co/datasets/laion/terminal-bench-2-1-apptainer-v1.text-generationn<1K1 likes1.8k downloads11d agoHugging Face22camel-ai /terminal-bench-core_migrated0 likes1.7k downloads6mo agoHugging Face23DCAgent /GPT-5-terminal-bench-2textn<1K0 likes1.4k downloads11mo agoHugging Face24wufeiwu /Terminal-Bench-Evo TerminalBench-Evo Data This directory contains the TerminalBench-Evo dataset, which is part of the EvoArena benchmark suite introduced in the paper EvoArena: Tracking Memory Evolution for Robust LLM Agents in Dynamic Environments. Project Page | GitHub Repository Dataset Description TerminalBench-Evo includes the original TerminalBench tasks and several corresponding EVO versions generated for each discovered task family in the repository. It focuses on executable… See the full description on the dataset page: https://huggingface.co/datasets/wufeiwu/Terminal-Bench-Evo.other1 likes1.2k downloads4mo agoHugging Face25camel-ai /terminal-bench-core-0.1.1_migrated0 likes1.2k downloads6mo agoHugging Face26JinhuaQIN /TerminalBenchV31 likes1.2k downloads20d agoHugging Face27YLR9933 /terminal-bench-science-trailtext10K<n<100K0 likes1.1k downloads3d agoHugging Face28camel-ai /terminal-bench-2.00 likes1.1k downloads6mo agoHugging Face29introvoyz041 /terminal-bench-2.0Warning: This is a read-only mirror. The primary source for this dataset is on GitHub: https://github.com/laude-institute/terminal-bench-2. Please open issues and pull requests there. How this mirror was created git clone https://github.com/laude-institute/terminal-bench-2.git /tmp/tb2-mirror cd /tmp/tb2-mirror git lfs install git lfs migrate import \ --include="*.tar.gz,*.png,*.jpg,*.jpeg,*.mp4,*.pt,*.pth,*.7z,*.wasm,*.gz,*.enc,train-fasttext/tests/private_test.txt" \ --everything git… See the full description on the dataset page: https://huggingface.co/datasets/introvoyz041/terminal-bench-2.0.text-generationn<1K0 likes931 downloads6mo agoHugging Face30openguardrails /terminal-bench-2.1-deepseek-v4-flash-trajectories DeepSeek-V4-Flash on Terminal-Bench 2.1 — full agent trajectories Every agent trajectory from a controlled scaffold comparison: the same model, the same machine, the same 89 tasks, only the agent harness changed. Scaffold Solved terminus-2 (Terminal-Bench's own agent) 53 / 89 dsh sdk-minimal (DeepSeek Harness) 61 / 89 Paired: dsh solved 18 tasks terminus-2 missed, terminus-2 solved 8 that dsh missed, 43 both, 18 neither, 2 not scorable (see Caveats).… See the full description on the dataset page: https://huggingface.co/datasets/openguardrails/terminal-bench-2.1-deepseek-v4-flash-trajectories.n<1K0 likes804 downloads24d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.