Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tasksource /bigbenchBIG-Bench but it doesn't require the hellish dependencies (tensorflow, pypi-bigbench, protobuf) of the official version. dataset = load_dataset("tasksource/bigbench",'movie_recommendation') Code to reproduce: https://colab.research.google.com/drive/1MKdLdF7oqrSQCeavAcsEnPdI85kD0LzU?usp=sharing Datasets are capped to 50k examples to keep things light. I also removed the default split when train was available also to save space, as default=train+val. @article{srivastava2022beyond… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/bigbench.textmultiple-choice100K<n<1M69 likes14k downloads1y agoHugging Face02facebook /kilt_tasks Dataset Card for KILT Dataset Summary KILT has been built from 11 datasets representing 5 types of tasks: Fact-checking Entity linking Slot filling Open domain QA Dialog generation All these datasets have been grounded in a single pre-processed Wikipedia dump, allowing for fairer and more consistent evaluation as well as enabling new task setups such as multitask and transfer learning with minimal effort. KILT also provides tools to analyze and understand the… See the full description on the dataset page: https://huggingface.co/datasets/facebook/kilt_tasks.textfill-mask1M<n<10M68 likes3.5k downloads3y agoHugging Face03tasksource /tasksource-instruct tasksource-instruct Instruction-tuning data recast from the ~480 English classification, multiple-choice and token-classification tasks of tasksource. Every example comes from a human-built dataset (NLI, logical reasoning, sentiment, hate speech, discourse, argumentation, ...), not from a teacher model. Each task is capped at 30k training examples, so no task dominates. Many tasks aren't in FLAN v2, for example DynaSent, DynaHate, discriminative bAbI, epistemic logic, RuleTaker… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/tasksource-instruct.texttext-generation1M<n<10M24 likes2.6k downloads13d agoHugging Face04FACET-Terminal /FACET-Terminal-Tasks-6k FACET-Terminal-Tasks-6k 6,020 execution-grounded tasks for terminal agents, coding agents, and executable workflow research 🌐 FACET Project Website 📄 FACET Paper 💻 FACET-Terminal GitHub Repository 🤗 FACET-Terminal Models & Data Dataset Overview FACET-Terminal-Tasks-6k contains 6,020 public-release-ready Harbor tasks produced by FACET. Each task is an executable environment rather than a standalone prompt: it includes a natural-language instruction… See the full description on the dataset page: https://huggingface.co/datasets/FACET-Terminal/FACET-Terminal-Tasks-6k.texttext-generation1K<n<10K1 likes2.4k downloads2mo agoHugging Face05Bohan22 /MLS-Bench-Tasks MLS-Bench Tasks MLS-Bench is a benchmark for machine learning science. Where most agent benchmarks reward engineering one fixed instance — clean the data, tune the pipeline, climb a leaderboard — MLS-Bench asks the harder question: can an AI agent propose a new component, loss, optimizer, or training procedure whose gain transfers across settings, seeds, datasets, and scales?The benchmark contains 140 tasks across 12 ML research domains. Each task fixes a research scaffold… See the full description on the dataset page: https://huggingface.co/datasets/Bohan22/MLS-Bench-Tasks.texttext-generationn<1K8 likes2k downloads5mo agoHugging Face06SegunOni /osworld_tasks_filestext-classification1M<n<10M0 likes1.9k downloads8mo agoHugging Face07yatin-superintelligence /Creative-Professionals-Agentic-Tasks-1M Creative Professionals Agentic Tasks (1M) Abstract A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Creative-Professionals-Agentic-Tasks-1M.tabulartext-generation1M<n<10M29 likes1.3k downloads7mo agoHugging Face08open-athena /pdbthink-coordinate-tasks PDBThink Coordinate Tasks 100,000 new coordinate-interpretation tasks across all 19 active PDBThink families, from 2,671 experimental PDB entries in 1,867 source groups. Version 1.3.0; deterministic seed 2026100101. The model receives sanitised, rotated, rounded protein coordinates and a question. It must answer without tools. This release contains no sequence-to-structure prediction tasks and no retired MECH tasks. It is intended for additional evaluation, RL with deterministic… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/pdbthink-coordinate-tasks.tabulartext-generation100K<n<1M0 likes1.2k downloads8d agoHugging Face09tasksource /FOL-nli Dataset Card for "FOL-nli" https://github.com/sileod/unigram/ https://arxiv.org/abs/2406.11035 Citation: @article{sileo2024scaling, title={Scaling Synthetic Logical Reasoning Datasets with Context-Sensitive Declarative Grammars}, author={Sileo, Damien}, journal={arXiv preprint arXiv:2406.11035}, year={2024} } texttext-classification100K<n<1M3 likes788 downloads9mo agoHugging Face10jablonkagroup /corral-environment-tasks Corral – Environment Tasks Task definitions across the 8 Corral environments, including descriptions, allowed tools, scoring functions, and submission formats 📋 Dataset Summary This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the task definitions for the 8 environments included in the Corral benchmark. The dataset is organized into multiple configurations… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral-environment-tasks.texttext-generationn<1K0 likes735 downloads4mo agoHugging Face11yatin-superintelligence /Audio-Video-Engineering-Agentic-Tasks-1M Audio/Video Engineering Agentic Tasks (1M) Abstract A highly specialized dataset comprising 1,029,459 in-context troubleshooting prompts and execution commands built for the deepest levels of media production. Unlike standard datasets that simulate clean, theoretical instructions, this matrix captures the chaotic, highly-detailed, and conversational reality of professional audio engineers, composers, and video editors mid-session. It is engineered to train multimodal AI… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Audio-Video-Engineering-Agentic-Tasks-1M.tabulartext-generation1M<n<10M14 likes599 downloads7mo agoHugging Face12rAVEUK /Creative-Professionals-Agentic-Tasks-1M Creative Professionals Agentic Tasks (1M) Abstract A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/rAVEUK/Creative-Professionals-Agentic-Tasks-1M.tabulartext-generation1M<n<10M3 likes554 downloads7mo agoHugging Face13muahmed7338 /kernelascent-tasks KernelAscent — public dev split KernelAscent is a benchmark for recursive self-improvement (RSI): a model optimizes the GPU kernels used to train itself, and we measure whether kernel-optimization capability compounds across rounds. This is the public dev split, released for self-benchmarking and research; the leaderboard is scored on a private held-out split. Project & code: https://github.com/ahmd-mohsin/KernelAscent Leaderboard & docs:… See the full description on the dataset page: https://huggingface.co/datasets/muahmed7338/kernelascent-tasks.text-generationn<1K0 likes548 downloads22d agoHugging Face14yzxjb /long-gui-tasks-v1 Long GUI Tasks HF Release v1 This directory is a local release package prepared for publishing the long-horizon GUI task dataset to Hugging Face. Contents metadata/train.jsonl: training split in ms-swift style messages + images format metadata/test.jsonl: test split in the same format metadata/train_manifest.jsonl: training sample manifest with metadata and image_ids metadata/test_manifest.jsonl: test sample manifest indexes/image_index.jsonl: unique image registry with… See the full description on the dataset page: https://huggingface.co/datasets/yzxjb/long-gui-tasks-v1.image-text-to-text10K<n<100K0 likes490 downloads6mo agoHugging Face15kryp1234 /Creative-Professionals-Agentic-Tasks-1M Creative Professionals Agentic Tasks (1M) Abstract A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/kryp1234/Creative-Professionals-Agentic-Tasks-1M.tabulartext-generation1M<n<10M1 likes474 downloads7mo agoHugging Face16whitecircle /swe-rebench-v2-clean-python-tasks SWE-rebench-V2 clean Python tasks A train/test split of Python tasks from nebius/SWE-rebench-V2. We took the Python subset of the original dataset and kept only the tasks where the golden patch passes the unit tests and the empty patch does not. train: 3,837 instances from 408 repositories test: 500 instances from 100 repositories The split is made by repository, so no repository appears in both splits. We evaluated multiple models on the test split as of June 2026 — the… See the full description on the dataset page: https://huggingface.co/datasets/whitecircle/swe-rebench-v2-clean-python-tasks.texttext-generation1K<n<10K1 likes362 downloads3mo agoHugging Face17Fzz1 /SWE-Rebench-Tasks-Clean SWE-Rebench-Tasks-Clean 1,317 verified-solvable, contamination-controlled software-engineering tasks for terminal-agent RL training. Adapted from nebius/SWE-rebench-V2 (real GitHub issue → PR tasks with executable test contracts) into the TerminalWorld task format. Companion dataset to Fzz1/SWE-Smith-Seeds-Clean, same layout. Every task is a directory containing: file content instruction.md the issue text the agent sees (plus linked issue discussion where available)… See the full description on the dataset page: https://huggingface.co/datasets/Fzz1/SWE-Rebench-Tasks-Clean.tabulartext-generation1K<n<10K1 likes359 downloads2mo agoHugging Face18anonymous-dataset-submission-warp /warp-taskgen-generated-ipi-tasks-50 WARP Taskgen Generated IPI Tasks 50 Dataset Summary This dataset contains WARP Taskgen Phase 4 browser-agent trajectories for a 50-task generated indirect prompt injection (IPI) cohort. The trajectories were produced with the AgentLab harness on WebArena GitLab and Postmill (Reddit) benchmark applications. The export is a report-only projection of already written benchmark artifacts. It does not alter scoring, PVPO encounter checks, rewards, admission, or trajectory… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-dataset-submission-warp/warp-taskgen-generated-ipi-tasks-50.text-generationn<1K0 likes340 downloads5mo agoHugging Face19tasksource /icl-symbol-tuning-instruct Description Few-shot prompting demonstrates that language models can learn in context even though they were not trained to do. However, explicitly learning to learn in context meta-icl leads to better results. With symbol tuning, labels are replaced with arbitrary symbols (e.g. foo/bar), which makes learning in context a key condition to learn the instructions We implement symbol tuning, as presented in the Symbol tuning improves in-context learning paper with tasksource… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/icl-symbol-tuning-instruct.texttext-classification100K<n<1M20 likes293 downloads3y agoHugging Face20tasksource /tasksource_dpo_pairs tasksource_dpo_pairs Preference pairs from ~470 human-annotated classification and multiple-choice tasks of tasksource, for DPO and other preference training. Each pair takes one example from tasksource-instruct. chosen is the gold answer, and rejected is another answer the prompt offers. No text is generated by a model: the preference comes from the source labels, notably on NLI, logical reasoning, and other expert-built tasks. from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/tasksource_dpo_pairs.texttext-generation1M<n<10M21 likes283 downloads13d agoHugging Face21tasksource /prm800k_dpo PRM800K preference pairs (prompt / chosen / rejected) PRM800K (OpenAI, Let's Verify Step by Step) turned into clean preference pairs, ready for DPO / reward modeling (TRL "standard" preference format). Raw source: tasksource/PRM800K. Train/test follow the original MATH benchmark splits (see Splits below). from datasets import load_dataset ds = load_dataset("tasksource/prm800k_dpo", "solution") # or "step" Configs solution — whole-solution preference… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/prm800k_dpo.texttext-generation10K<n<100K4 likes232 downloads16d agoHugging Face22ServiceNow-AI /servicenow-tasks ServiceNow Tasks An evaluation dataset for web agents operating on ServiceNow instances. Each row contains a task goal, a configuration dict, and a standalone Python validation function that scores agent performance by querying the ServiceNow API. Derived from the WorkArena L1 benchmark. Dataset overview 330 samples across 33 task types grouped into 8 categories Each task has 10 seeded variants Validators are self-contained Python functions that call the… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow-AI/servicenow-tasks.texttext-generationn<1K0 likes198 downloads5mo agoHugging Face23Emulated-Inc /hard-reasoning-tasks-training-pool Hard reasoning tasks training pool Reasoning questions in 22 families, each with a short exact answer, so a trained model can be marked against the key by string comparison and no judge is needed. Part of it was drawn here by program under the seed named below, part of it was read from two public repositories at the commits named below. The whole is laid out twice, once rewritten into a single shape and once as its sources publish it. Train on either layer or on both.… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/hard-reasoning-tasks-training-pool.textquestion-answering100K<n<1M0 likes188 downloads28d agoHugging Face24osieosie /tmax-tasks-selfgen-qwen35-9b-20260919-1k-verified tmax self-generated tasks — Qwen3.5-9B (verified arm) The 1,042 tasks from …-20260919-1k, each graded against the issue-#12 rubric by the same model that generated them (hamishivi/Qwen3.5-9B). Using the generator as its own reviewer is deliberate: the question is whether an open-weights model can carry both halves of the loop. A stronger reviewer would answer a different question. The grader sees instruction / setup.sh / tests only. truth is withheld from it, so it is no better… See the full description on the dataset page: https://huggingface.co/datasets/osieosie/tmax-tasks-selfgen-qwen35-9b-20260919-1k-verified.tabulartext-generation1K<n<10K0 likes187 downloads20d agoHugging Face25vrushankpatel5 /taskstreamTaskStream is a comprehensive dataset of enterprise business workflows, decision-making processes, organizational structures, and operational documentation across multiple industries.text-generation1K<n<10K2 likes176 downloads1y agoHugging Face26osieosie /tmax-tasks-selfgen-qwen35-9b-20260919-1k tmax self-generated tasks — Qwen3.5-9B (raw arm) 1,042 agentic terminal tasks generated by hamishivi/Qwen3.5-9B using the rl_data pipeline. Every previous tmax corpus was written by an API model (gemini-3.1-pro-preview); this one asks whether the model we RL on can generate its own training data. Companion: …-1k-verified — the same 1,042 tasks with rubric labels from the same model acting as reviewer. Recipe Identical to the Gemini 1k run… See the full description on the dataset page: https://huggingface.co/datasets/osieosie/tmax-tasks-selfgen-qwen35-9b-20260919-1k.texttext-generation1K<n<10K0 likes166 downloads21d agoHugging Face27joshuaswarren /h6-failure-gate-tasks H6 Failure-Gate Trap Tasks 30 synthetic TypeScript repair tasks built to measure whether an LLM coding agent repeats a known, previously observed failure. Each task contains a deliberate "trap": a wrong fix that looks more attractive than the correct one. The dataset was built for a preregistered experiment on failure-memory delivery timing (paper link will be added on publication; study registration and harness: https://github.com/joshuaswarren/remnic/issues/1963). canary GUID:… See the full description on the dataset page: https://huggingface.co/datasets/joshuaswarren/h6-failure-gate-tasks.text-generationn<1K0 likes163 downloads2mo agoHugging Face28kshitijthakkar /nirnaya-jevstyle-tasksource-v2 Nirnaya Typed Decisions ChatML v2 A provenance-preserving ChatML conversion of Apache-2.0 typed-decision training data. Each conversation keeps the original state, question(s), and typed answer(s): state is the system message, questions are the user message, and answers are the assistant message. No synthetic instruction wrapper is added. Sources and exact inclusion rules Source Revision Config / split Filter Input retained Output ChatML rows… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/nirnaya-jevstyle-tasksource-v2.texttext-generation100K<n<1M0 likes116 downloads5d agoHugging Face29asingh15 /rubric-diversity-tasks Rubric Diversity Tasks The default qwen_core configuration contains 1,296,719 curated task groups after screening 108,713 foundational tasks into the separate foundational configuration. The tasks configuration retains all 1,405,432 curated task groups. Difficulty is labeled for a Qwen3.6-27B curriculum using transparent source and structural priors. No new Qwen3.6-27B success rates were measured. core_candidate means an explicit nonfoundational source/structure signal;… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/rubric-diversity-tasks.tabulartext-generation1M<n<10M0 likes114 downloads12d agoHugging Face301-800-SHARED-TASKS /xlsum-subset Dataset Card for "XL-Sum" Dataset Summary We present XLSum, a comprehensive and diverse dataset comprising 1.35 million professionally annotated article-summary pairs from BBC, extracted using a set of carefully designed heuristics. The dataset covers 45 languages ranging from low to high-resource, for many of which no public dataset is currently available. XL-Sum is highly abstractive, concise, and of high quality, as indicated by human and intrinsic evaluation.… See the full description on the dataset page: https://huggingface.co/datasets/1-800-SHARED-TASKS/xlsum-subset.textsummarizationn<1K0 likes111 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.