datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bigbenchBIG-Bench but it doesn't require the hellish dependencies (tensorflow, pypi-bigbench, protobuf) of the official version.
dataset = load_dataset("tasksource/bigbench",'movie_recommendation')
Code to reproduce:
https://colab.research.google.com/drive/1MKdLdF7oqrSQCeavAcsEnPdI85kD0LzU?usp=sharing
Datasets are capped to 50k examples to keep things light.
I also removed the default split when train was available also to save space, as default=train+val.
@article{srivastava2022beyond… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/bigbench.kilt_tasks
Dataset Card for KILT
Dataset Summary
KILT has been built from 11 datasets representing 5 types of tasks:
Fact-checking
Entity linking
Slot filling
Open domain QA
Dialog generation
All these datasets have been grounded in a single pre-processed Wikipedia dump, allowing for fairer and more consistent evaluation as well as enabling new task setups such as multitask and transfer learning with minimal effort. KILT also provides tools to analyze and understand the… See the full description on the dataset page: https://huggingface.co/datasets/facebook/kilt_tasks.tasksource-instruct
tasksource-instruct
Instruction-tuning data recast from the ~480 English classification, multiple-choice
and token-classification tasks of tasksource.
Every example comes from a human-built dataset (NLI, logical reasoning, sentiment,
hate speech, discourse, argumentation, ...), not from a teacher model. Each task is
capped at 30k training examples, so no task dominates. Many tasks aren't in FLAN v2,
for example DynaSent, DynaHate, discriminative bAbI, epistemic logic, RuleTaker… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/tasksource-instruct.FACET-Terminal-Tasks-6k
FACET-Terminal-Tasks-6k
6,020 execution-grounded tasks for terminal agents, coding agents, and executable workflow research
🌐 FACET Project Website
📄 FACET Paper
💻 FACET-Terminal GitHub Repository
🤗 FACET-Terminal Models & Data
Dataset Overview
FACET-Terminal-Tasks-6k contains 6,020 public-release-ready Harbor tasks produced by FACET. Each task is an executable environment rather than a standalone prompt: it includes a natural-language instruction… See the full description on the dataset page: https://huggingface.co/datasets/FACET-Terminal/FACET-Terminal-Tasks-6k.MLS-Bench-Tasks
MLS-Bench Tasks
MLS-Bench is a benchmark for machine learning science. Where most agent benchmarks reward engineering one fixed instance — clean the data, tune the pipeline, climb a leaderboard — MLS-Bench asks the harder question: can an AI agent propose a new component, loss, optimizer, or training procedure whose gain transfers across settings, seeds, datasets, and scales?The benchmark contains 140 tasks across 12 ML research domains. Each task fixes a research scaffold… See the full description on the dataset page: https://huggingface.co/datasets/Bohan22/MLS-Bench-Tasks.osworld_tasks_filesCreative-Professionals-Agentic-Tasks-1M
Creative Professionals Agentic Tasks (1M)
Abstract
A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Creative-Professionals-Agentic-Tasks-1M.pdbthink-coordinate-tasks
PDBThink Coordinate Tasks
100,000 new coordinate-interpretation tasks across all 19 active PDBThink families,
from 2,671 experimental PDB entries in 1,867 source groups.
Version 1.3.0; deterministic seed 2026100101.
The model receives sanitised, rotated, rounded protein coordinates and a question.
It must answer without tools. This release contains no sequence-to-structure
prediction tasks and no retired MECH tasks. It is intended for additional evaluation,
RL with deterministic… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/pdbthink-coordinate-tasks.FOL-nli
Dataset Card for "FOL-nli"
https://github.com/sileod/unigram/
https://arxiv.org/abs/2406.11035
Citation:
@article{sileo2024scaling,
title={Scaling Synthetic Logical Reasoning Datasets with Context-Sensitive Declarative Grammars},
author={Sileo, Damien},
journal={arXiv preprint arXiv:2406.11035},
year={2024}
}
corral-environment-tasks
Corral – Environment Tasks
Task definitions across the 8 Corral environments, including descriptions, allowed tools, scoring functions, and submission formats
📋 Dataset Summary
This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the task definitions for the 8 environments included in the Corral benchmark.
The dataset is organized into multiple configurations… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral-environment-tasks.Audio-Video-Engineering-Agentic-Tasks-1M
Audio/Video Engineering Agentic Tasks (1M)
Abstract
A highly specialized dataset comprising 1,029,459 in-context troubleshooting prompts and execution commands built for the deepest levels of media production. Unlike standard datasets that simulate clean, theoretical instructions, this matrix captures the chaotic, highly-detailed, and conversational reality of professional audio engineers, composers, and video editors mid-session. It is engineered to train multimodal AI… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Audio-Video-Engineering-Agentic-Tasks-1M.Creative-Professionals-Agentic-Tasks-1M
Creative Professionals Agentic Tasks (1M)
Abstract
A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/rAVEUK/Creative-Professionals-Agentic-Tasks-1M.kernelascent-tasks
KernelAscent — public dev split
KernelAscent is a benchmark for recursive self-improvement (RSI): a model optimizes
the GPU kernels used to train itself, and we measure whether kernel-optimization
capability compounds across rounds. This is the public dev split, released for
self-benchmarking and research; the leaderboard is scored on a private held-out split.
Project & code: https://github.com/ahmd-mohsin/KernelAscent
Leaderboard & docs:… See the full description on the dataset page: https://huggingface.co/datasets/muahmed7338/kernelascent-tasks.long-gui-tasks-v1
Long GUI Tasks HF Release v1
This directory is a local release package prepared for publishing the long-horizon GUI task dataset to Hugging Face.
Contents
metadata/train.jsonl: training split in ms-swift style messages + images format
metadata/test.jsonl: test split in the same format
metadata/train_manifest.jsonl: training sample manifest with metadata and image_ids
metadata/test_manifest.jsonl: test sample manifest
indexes/image_index.jsonl: unique image registry with… See the full description on the dataset page: https://huggingface.co/datasets/yzxjb/long-gui-tasks-v1.Creative-Professionals-Agentic-Tasks-1M
Creative Professionals Agentic Tasks (1M)
Abstract
A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/kryp1234/Creative-Professionals-Agentic-Tasks-1M.swe-rebench-v2-clean-python-tasks
SWE-rebench-V2 clean Python tasks
A train/test split of Python tasks from nebius/SWE-rebench-V2.
We took the Python subset of the original dataset and kept only the tasks where the golden patch passes
the unit tests and the empty patch does not.
train: 3,837 instances from 408 repositories
test: 500 instances from 100 repositories
The split is made by repository, so no repository appears in both splits.
We evaluated multiple models on the test split as of June 2026 — the… See the full description on the dataset page: https://huggingface.co/datasets/whitecircle/swe-rebench-v2-clean-python-tasks.SWE-Rebench-Tasks-Clean
SWE-Rebench-Tasks-Clean
1,317 verified-solvable, contamination-controlled software-engineering tasks for terminal-agent RL training.
Adapted from nebius/SWE-rebench-V2 (real GitHub issue → PR tasks with executable test contracts) into the TerminalWorld task format. Companion dataset to Fzz1/SWE-Smith-Seeds-Clean, same layout.
Every task is a directory containing:
file
content
instruction.md
the issue text the agent sees (plus linked issue discussion where available)… See the full description on the dataset page: https://huggingface.co/datasets/Fzz1/SWE-Rebench-Tasks-Clean.warp-taskgen-generated-ipi-tasks-50
WARP Taskgen Generated IPI Tasks 50
Dataset Summary
This dataset contains WARP Taskgen Phase 4 browser-agent trajectories for a
50-task generated indirect prompt injection (IPI) cohort. The trajectories were
produced with the AgentLab harness on
WebArena GitLab and Postmill (Reddit) benchmark applications.
The export is a report-only projection of already written benchmark artifacts.
It does not alter scoring, PVPO encounter checks, rewards, admission, or
trajectory… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-dataset-submission-warp/warp-taskgen-generated-ipi-tasks-50.icl-symbol-tuning-instruct
Description
Few-shot prompting demonstrates that language models can learn in context even though they were not trained to do. However, explicitly learning to learn in context meta-icl leads to better results. With symbol tuning, labels are replaced with arbitrary symbols (e.g. foo/bar), which makes learning in context a key condition to learn the instructions
We implement symbol tuning, as presented in the Symbol tuning improves in-context learning paper with tasksource… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/icl-symbol-tuning-instruct.tasksource_dpo_pairs
tasksource_dpo_pairs
Preference pairs from ~470 human-annotated classification and multiple-choice tasks
of tasksource, for DPO and other preference training.
Each pair takes one example from
tasksource-instruct.
chosen is the gold answer, and rejected is another answer the prompt offers.
No text is generated by a model: the preference comes from the source labels, notably
on NLI, logical reasoning, and other expert-built tasks.
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/tasksource_dpo_pairs.prm800k_dpo
PRM800K preference pairs (prompt / chosen / rejected)
PRM800K (OpenAI, Let's Verify Step by Step) turned into clean
preference pairs, ready for DPO / reward modeling (TRL "standard" preference format).
Raw source: tasksource/PRM800K.
Train/test follow the original MATH benchmark splits (see Splits below).
from datasets import load_dataset
ds = load_dataset("tasksource/prm800k_dpo", "solution") # or "step"
Configs
solution — whole-solution preference… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/prm800k_dpo.servicenow-tasks
ServiceNow Tasks
An evaluation dataset for web agents operating on ServiceNow instances. Each row contains a task goal, a configuration dict, and a standalone Python validation function that scores agent performance by querying the ServiceNow API.
Derived from the WorkArena L1 benchmark.
Dataset overview
330 samples across 33 task types grouped into 8 categories
Each task has 10 seeded variants
Validators are self-contained Python functions that call the… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow-AI/servicenow-tasks.hard-reasoning-tasks-training-pool
Hard reasoning tasks training pool
Reasoning questions in 22 families, each with a short exact answer, so a trained model can be
marked against the key by string comparison and no judge is needed. Part of it was drawn here by
program under the seed named below, part of it was read from two public repositories at the
commits named below. The whole is laid out twice, once rewritten into a single shape and once as
its sources publish it. Train on either layer or on both.… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/hard-reasoning-tasks-training-pool.tmax-tasks-selfgen-qwen35-9b-20260919-1k-verified
tmax self-generated tasks — Qwen3.5-9B (verified arm)
The 1,042 tasks from
…-20260919-1k,
each graded against the issue-#12 rubric by the same model that generated them
(hamishivi/Qwen3.5-9B). Using the generator as its own reviewer is deliberate: the
question is whether an open-weights model can carry both halves of the loop. A stronger
reviewer would answer a different question.
The grader sees instruction / setup.sh / tests only. truth is withheld from it,
so it is no better… See the full description on the dataset page: https://huggingface.co/datasets/osieosie/tmax-tasks-selfgen-qwen35-9b-20260919-1k-verified.taskstreamTaskStream is a comprehensive dataset of enterprise business workflows, decision-making processes,
organizational structures, and operational documentation across multiple industries.tmax-tasks-selfgen-qwen35-9b-20260919-1k
tmax self-generated tasks — Qwen3.5-9B (raw arm)
1,042 agentic terminal tasks generated by hamishivi/Qwen3.5-9B using the rl_data
pipeline. Every previous tmax corpus was written by an API model (gemini-3.1-pro-preview);
this one asks whether the model we RL on can generate its own training data.
Companion: …-1k-verified
— the same 1,042 tasks with rubric labels from the same model acting as reviewer.
Recipe
Identical to the Gemini 1k run… See the full description on the dataset page: https://huggingface.co/datasets/osieosie/tmax-tasks-selfgen-qwen35-9b-20260919-1k.h6-failure-gate-tasks
H6 Failure-Gate Trap Tasks
30 synthetic TypeScript repair tasks built to measure whether an LLM coding
agent repeats a known, previously observed failure. Each task contains a
deliberate "trap": a wrong fix that looks more attractive than the correct
one. The dataset was built for a preregistered experiment on failure-memory
delivery timing (paper link will be added on publication; study registration
and harness: https://github.com/joshuaswarren/remnic/issues/1963).
canary GUID:… See the full description on the dataset page: https://huggingface.co/datasets/joshuaswarren/h6-failure-gate-tasks.nirnaya-jevstyle-tasksource-v2
Nirnaya Typed Decisions ChatML v2
A provenance-preserving ChatML conversion of Apache-2.0 typed-decision training data.
Each conversation keeps the original state, question(s), and typed answer(s): state is the system message, questions are the user message, and answers are the assistant message. No synthetic instruction wrapper is added.
Sources and exact inclusion rules
Source
Revision
Config / split
Filter
Input retained
Output ChatML rows… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/nirnaya-jevstyle-tasksource-v2.rubric-diversity-tasks
Rubric Diversity Tasks
The default qwen_core configuration contains 1,296,719 curated task groups after screening 108,713 foundational tasks into the separate foundational configuration. The tasks configuration retains all 1,405,432 curated task groups.
Difficulty is labeled for a Qwen3.6-27B curriculum using transparent source and structural priors. No new Qwen3.6-27B success rates were measured. core_candidate means an explicit nonfoundational source/structure signal;… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/rubric-diversity-tasks.xlsum-subset
Dataset Card for "XL-Sum"
Dataset Summary
We present XLSum, a comprehensive and diverse dataset comprising 1.35 million professionally annotated article-summary pairs from BBC, extracted using a set of carefully designed heuristics. The dataset covers 45 languages ranging from low to high-resource, for many of which no public dataset is currently available. XL-Sum is highly abstractive, concise, and of high quality, as indicated by human and intrinsic evaluation.… See the full description on the dataset page: https://huggingface.co/datasets/1-800-SHARED-TASKS/xlsum-subset.
