tasksource/tasksource-instruct
tasksource-instruct Instruction-tuning data recast from the ~480 English classification, multiple-choice and token-classification tasks of tasksource. Every example comes from a human-built dataset (NLI, logical reasoning, sentiment, hate speech, discourse, argumentation, ...), not from a teacher model. Each task is capped at 30k training examples, so no task dominates. Many tasks aren't in FLAN v2, for example DynaSent, DynaHate, discriminative bAbI, epistemic logic, RuleTaker… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/tasksource-instruct.
tasksource-instruct
Instruction-tuning data recast from the ~480 English classification, multiple-choice and token-classification tasks of [tasksource](https://github.com/sileod/tasksource).
Every example comes from a human-built dataset (NLI, logical reasoning, sentiment, hate speech, discourse, argumentation, ...), not from a teacher model. Each task is capped at 30k training examples, so no task dominates. Many tasks aren't in FLAN v2, for example DynaSent, DynaHate, discriminative bAbI, epistemic logic, RuleTaker, veridicality and dozens of NLI datasets.
from datasets import load_dataset
ds = load_dataset("tasksource/tasksource-instruct", split="train")
ds = ds.filter(lambda use: use == "commercial", input_columns="license_use") # optionalFormat
Prompts ask for the answer with no explanation, so the short targets don't teach a model to stop explaining in general. Tasks are interleaved round-robin, so any slice of the split mixes them. Validation and test keep up to 500 examples per task.
tasksource-instruct works well mixed with FLAN v2 or other instruction data. It covers discriminative reasoning tasks that those sets cover less.
For preference pairs built from the same rows, see tasksource_dpo_pairs. For soft labels, ratings and multi-question requests, see tasksource-jev-typed-decisions.
Reproducibility
The dataset is built by `scripts/build_instruct_dataset.py`:
PYTHONPATH=.:src python scripts/build_instruct_dataset.py --finalizeSources are loaded at pinned Hub revisions. sources.yaml records, per task, the Hub dataset, revision, original dataset, licenses and row counts, and build-report.jsonl records the code commit of each task's build. MMLU, BIG-bench and BLiMP are left out, so they stay clean for evaluation. Other public benchmarks (GLUE, SuperGLUE, HellaSwag, PIQA, ...) are in the data through their training splits.
License and scope
Tasksource harmonizes datasets from many publishers; their original licenses and terms still apply, hence license: other.
licenselists thelicenseof the Hub dataset card the task was loaded from, and of the original dataset behind a tasksource copy. It also lists licenses recorded by the Data Provenance Initiative, marked(DPI).license_usetakes the most restrictive of those:non-commercialif any is non-commercial or academic-only,commercialif one allows commercial use (share-alike and copyleft included), andunspecifiedotherwise.
This is a best-effort aid, not legal advice. Check the original terms before relying on them.
Citation
@inproceedings{sileo-2024-tasksource,
title = "tasksource: A Large Collection of {NLP} tasks with a Structured Dataset Preprocessing Framework",
author = "Sileo, Damien",
booktitle = "Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)",
month = may,
year = "2024",
address = "Torino, Italia",
publisher = "ELRA and ICCL",
url = "https://aclanthology.org/2024.lrec-main.1361/",
pages = "15655--15684",
}