Team Ai
Datasetpublic

LocalLLaMA/terminal-bench-mini

terminal-bench-mini Fourteen of Terminal-Bench 2.0's ninety tasks, picked so that ranking agents on the subset reproduces ranking them on the whole benchmark. Running ninety tasks five times each is how the official leaderboard is built. That is out of reach if you are comparing quant variants, fine-tunes or local models on your own hardware. This subset turns a multi-day sweep into a few hours. Same approach as deepswe-mini: take the published per-task results, rank the field… See the full description on the dataset page: https://huggingface.co/datasets/LocalLLaMA/terminal-bench-mini.

sourceHugging Faceapache-2.0updated 21d agoView on Hugging Face
6likes2.9kdownloads

LocalLLaMA/terminal-bench-mini · main · files are served by the source, never re-hosted here