datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
terminal-bench-2.1
Terminal-Bench 2.1 (Harbor git-repos dataset)
Harbor website · Harbor GitHub
This is a private mirror of the task content from
harbor-framework/terminal-bench-2-1
at commit 7131e43
(the source repo has no tagged releases yet), laid out so it can be consumed directly
by Harbor's
git-repos dataset support.
The primary source is the GitHub repository above — please open issues and pull
requests there, not here.
How to run
Always pass the full URL, not org/name — a… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/terminal-bench-2.1.terminal-bench-3.0
Terminal-Bench 3.0
The primary source is hosted on GitHub, please open issues and pull
requests there, not here.
The official published dataset is hosted on the Harbor Hub along with the official leaderboard. Usage e.g. harbor run -d terminal-bench/terminal-bench@3.0.0
This repo is a mirror of harbor-framework/terminal-bench
at tag v3.0.0, laid out so it can be consumed directly by
Harbor's
git-repos dataset support.
How to run via this Huggingface repo
Always… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/terminal-bench-3.0.terminal-bench
Terminal-Bench
The primary source is hosted on GitHub, please open issues and
pull requests there, not here.
Terminal-Bench is now a continuous benchmark: new versions are released periodically as tags on the source repo
instead of one-off snapshots. This dataset mirrors that model on the Hub: instead of a separate
terminal-bench-X.Y repo per release, one repo, tagged per version. main always tracks the latest published
version; each release is additionally available as an… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/terminal-bench.terminal-bench-2.0Warning: The leaderboard above is unofficial. The official leaderboard is https://www.tbench.ai/leaderboard/terminal-bench/2.0, in which entires are audited for correct configuration, results show which agent harness is used, and verified trajectories are publicly viewable.
Warning: The dataset is a read-only mirror. The primary source for this dataset is on GitHub: https://github.com/harbor-framework/terminal-bench-2. Please open issues and pull requests there.
How this mirror was created… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/terminal-bench-2.0.terminal-bench-science
Terminal-Bench-Science
The primary source is hosted on GitHub, please open
issues and pull requests there, not here.
Terminal-Bench-Science is a benchmark of real-world computational research
workflows across the life, physical, earth, mathematical, and engineering sciences. Like Terminal-Bench, it's a
continuous benchmark: releases are published as tags on the source repo. This dataset mirrors that on the Hub:
one repo, tagged per version, instead of a separate repo per… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/terminal-bench-science.terminal-bench-science-lfs
Terminal-Bench-Science — task input mirror
Large input files for Terminal-Bench-Science
tasks, which cannot be committed to git. Tasks pull from here at container build
time, pinned to a commit SHA and verified against a checksum file that ships in
the task directory.
One top-level prefix per task; everything lives under <task-name>/input/.
Benchmark contamination canary
This dataset is benchmark material. If you are assembling a training corpus,
exclude it.… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/terminal-bench-science-lfs.terminal-bench-2
Terminal-Bench-2.0 Beta
Welcome to Terminal-Bench-2.0! If you’re reading this you’re a member of the Terminal-Bench community that we’ve selected to get a sneak peek at the latest version of the benchmark.
Getting Started
First, clone Harbor (formerly “Sandboxes”):
git clone https://github.com/laude-institute/harbor.git
From inside the Harbor directory run:
uv sync
This will install Harbor, our new package for running agent evals.
You should now be able to run TB 2.0!… See the full description on the dataset page: https://huggingface.co/datasets/penfever/terminal-bench-2.terminal-bench-2-verified
Terminal-Bench 2.0 Verified: Instruction & Environment Fix Version
中文版本
2026.08.18 Update: Following the release of Terminal-Bench 2.1, we conducted another round of review on several tasks' grading and instructions, validating each fix end-to-end inside the actual task images. This round fixes 6 tasks in two categories: Grading/Test Fixes (dna-insert, make-doom-for-mips, filter-js-from-html, install-windows-3.11) — the test logic itself misjudged valid submissions or let… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/terminal-bench-2-verified.Long-Horizon-Terminal-Bench
Long-Horizon Terminal-Bench (LHTB)
LHTB is a 46-task benchmark for measuring how well LLM agents sustain useful
work in a containerized terminal over hundreds of steps. Unlike short-horizon
coding benchmarks where an agent writes one artifact and stops, LHTB drops the agent
into a stateful environment and grades it with hidden, rebuild-from-artifact
verifiers — self-reported progress does not count.
📝 Blog: https://zli12321.github.io/LHTB/
🏆 Leaderboard:… See the full description on the dataset page: https://huggingface.co/datasets/IntelligenceLab/Long-Horizon-Terminal-Bench.terminal-bench-2-leaderboard
Terminal-Bench 2.0 Leaderboard Submissions
This repository accepts leaderboard submissions for Terminal-Bench 2.0.
How to Submit
Fork this repository
Create a new branch for your submission
Add your submission (a job or folder of jobs) under submissions/terminal-bench/2.0/<agent>__<model(s)>/
Open a Pull Request
Submission Structure
submissions/
terminal-bench/
2.0/
<agent>__<model>/
metadata.yaml # Required: agent and model info… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/terminal-bench-2-leaderboard.terminal_bench_2
Terminal-Bench 2.0
######################################################################
# _____ _ _ ______________ #
# |_ _|__ _ __ _ __ ___ (_)_ __ __ _| | || || #
# | |/ _ \ '__| '_ ` _ \| | '_ \ / _` | | || > || #
# | | __/ | | | | | | | | | | | (_| | | || || #
# |_|\___|_| |_| |_| |_|_|_| |_|\__,_|_| ||____________|| #
# ____ _ ____… See the full description on the dataset page: https://huggingface.co/datasets/DCAgent2/terminal_bench_2.terminal-bench-lfsterminal-bench-pro
Terminal-Bench Pro
Overview
Terminal-Bench Pro is a systematic extension of the original Terminal-Bench, designed to address key limitations in existing terminal-agent benchmarks.
400 tasks (200 public + 200 private) across 8 domains: data processing, games, debugging, system admin, scientific computing, software engineering, ML, and security
Expert-designed tasks derived from real-world scenarios and GitHub issues
High test coverage with ~28.3 test cases per… See the full description on the dataset page: https://huggingface.co/datasets/alibabagroup/terminal-bench-pro.terminal-bench
Terminal-Bench Dataset
This dataset contains tasks from Terminal-Bench, a benchmark for evaluating AI agents in real terminal environments. Each task is packaged as a complete, self-contained archive that preserves the exact directory structure, binary files, Docker configurations, and test scripts needed for faithful reproduction.
The archive column contains a gzipped tarball of the entire task directory.
Dataset Overview
Terminal-Bench evaluates AI agents on… See the full description on the dataset page: https://huggingface.co/datasets/ia03/terminal-bench.terminalbench-trajectories
Terminal-Bench 2.0 Trajectories
Full agent trajectories from Terminal-Bench 2.0, a benchmark that evaluates AI coding agents on real-world terminal tasks. Each row is one trial: an agent attempting a task, with the complete step-by-step trace of messages, tool calls, and observations.
Explorer: yoonholee.com/web-apps/terminal-bench
Quick start
from datasets import load_dataset
import json
ds = load_dataset("yoonholee/terminalbench-trajectories", split="train")
#… See the full description on the dataset page: https://huggingface.co/datasets/yoonholee/terminalbench-trajectories.Terminal-Bench-Hard
Terminal-Bench Hard
Terminal-Bench Hard is a set of 100 challenging terminal-based agent tasks.
The tasks cover software engineering, debugging, data processing, system
administration, security, scientific computing, and related command-line
workflows.
Contents
tasks/: runnable tasks in Harbor format.
metadata/tasks.parquet: searchable task metadata and instructions.
Each task directory contains task.toml, instruction.md, an
environment/ directory, and verifier… See the full description on the dataset page: https://huggingface.co/datasets/Zhongzhi1228/Terminal-Bench-Hard.terminal-bench-mini
terminal-bench-mini
Fourteen of Terminal-Bench 2.0's ninety tasks, picked so that ranking agents on
the subset reproduces ranking them on the whole benchmark.
Running ninety tasks five times each is how the official leaderboard is built.
That is out of reach if you are comparing quant variants, fine-tunes or local
models on your own hardware. This subset turns a multi-day sweep into a few
hours.
Same approach as deepswe-mini:
take the published per-task results, rank the field… See the full description on the dataset page: https://huggingface.co/datasets/LocalLLaMA/terminal-bench-mini.terminal-bench-traces-localterminal-bench-2.1-qwen3.8-27b-traces
Terminal-Bench 2.1 traces: Qwen3.8-27B-GPTQ-4bit, xhigh / medium / low / off
Complete agent trajectories, verifier output, timing and token usage for all 89
Terminal-Bench 2.1 tasks run locally with
btbtyler09/Qwen3.8-27B-GPTQ-4bit
on 2× RTX 3090, plus the adaptive fallback reruns at lower reasoning effort.
Headline result: 62/89 (69.66%) at xhigh in a single clean pass.
Cumulative best-of across xhigh → medium → low → off fallbacks: 70/89 (78.65%).
The second number is not a… See the full description on the dataset page: https://huggingface.co/datasets/Lottolabs/terminal-bench-2.1-qwen3.8-27b-traces.terminalbench-sqlite-dbterminal-bench-2-1-apptainer-v1
Terminal-Bench 2.1 (offline Apptainer, v1)
A validated subset of Terminal-Bench 2.1 (revision 7131e4375048a0e408a8fb404b5f499d726b695b, Apache-2.0) for running on HPC clusters without Docker and without internet on compute nodes, with the harbor Apptainer bridge. 68 of the included tasks are byte-identical to upstream; 6 carry local changes (see below). Scores on this subset are not comparable to full 89-task TB2.1 leaderboard numbers.
status
tasks
meaning
validated… See the full description on the dataset page: https://huggingface.co/datasets/laion/terminal-bench-2-1-apptainer-v1.terminal-bench-core_migratedGPT-5-terminal-bench-2Terminal-Bench-Evo
TerminalBench-Evo Data
This directory contains the TerminalBench-Evo dataset, which is part of the EvoArena benchmark suite introduced in the paper EvoArena: Tracking Memory Evolution for Robust LLM Agents in Dynamic Environments.
Project Page | GitHub Repository
Dataset Description
TerminalBench-Evo includes the original TerminalBench tasks and several corresponding EVO versions generated for each discovered task family in the repository. It focuses on executable… See the full description on the dataset page: https://huggingface.co/datasets/wufeiwu/Terminal-Bench-Evo.terminal-bench-core-0.1.1_migratedTerminalBenchV3terminal-bench-science-trailterminal-bench-2.0terminal-bench-2.0Warning: This is a read-only mirror. The primary source for this dataset is on GitHub: https://github.com/laude-institute/terminal-bench-2. Please open issues and pull requests there.
How this mirror was created
git clone https://github.com/laude-institute/terminal-bench-2.git /tmp/tb2-mirror
cd /tmp/tb2-mirror
git lfs install
git lfs migrate import \
--include="*.tar.gz,*.png,*.jpg,*.jpeg,*.mp4,*.pt,*.pth,*.7z,*.wasm,*.gz,*.enc,train-fasttext/tests/private_test.txt" \
--everything
git… See the full description on the dataset page: https://huggingface.co/datasets/introvoyz041/terminal-bench-2.0.terminal-bench-2.1-deepseek-v4-flash-trajectories
DeepSeek-V4-Flash on Terminal-Bench 2.1 — full agent trajectories
Every agent trajectory from a controlled scaffold comparison: the same model, the
same machine, the same 89 tasks, only the agent harness changed.
Scaffold
Solved
terminus-2 (Terminal-Bench's own agent)
53 / 89
dsh sdk-minimal (DeepSeek Harness)
61 / 89
Paired: dsh solved 18 tasks terminus-2 missed, terminus-2 solved 8 that dsh
missed, 43 both, 18 neither, 2 not scorable (see Caveats).… See the full description on the dataset page: https://huggingface.co/datasets/openguardrails/terminal-bench-2.1-deepseek-v4-flash-trajectories.
