tasksource/tasksource-instruct
tasksource-instruct Instruction-tuning data recast from the ~480 English classification, multiple-choice and token-classification tasks of tasksource. Every example comes from a human-built dataset (NLI, logical reasoning, sentiment, hate speech, discourse, argumentation, ...), not from a teacher model. Each task is capped at 30k training examples, so no task dominates. Many tasks aren't in FLAN v2, for example DynaSent, DynaHate, discriminative bAbI, epistemic logic, RuleTaker… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/tasksource-instruct.
242.6k
1---2pretty_name: tasksource-instruct3language:4- en5license: other6size_categories:7- 1M<n<10M8task_categories:9- text-generation10- text-classification11- token-classification12- zero-shot-classification13tags:14- instructions15- instruction-tuning16- instruction-finetuning17- flan18- promptsource19- tasksource20configs:21- config_name: default22 data_files:23 - split: train24 path: data/train-*25 - split: test26 path: data/test-*27 - split: validation28 path: data/validation-*29dataset_info:30 features:31 - name: inputs32 dtype: string33 - name: targets34 dtype: string35 - name: task36 dtype: string37 - name: license38 dtype: string39 - name: license_use40 dtype: string41 splits:42 - name: train43 num_bytes: 650981849144 num_examples: 728225545 - name: validation46 num_bytes: 18388851547 num_examples: 18491448 - name: test49 num_bytes: 18411635250 num_examples: 19283051 download_size: 312279200752 dataset_size: 687782335853---54 55# tasksource-instruct56 57**Instruction-tuning data recast from the ~480 English classification, multiple-choice58and token-classification tasks of [tasksource](https://github.com/sileod/tasksource).**59 60Every example comes from a human-built dataset (NLI, logical reasoning, sentiment,61hate speech, discourse, argumentation, ...), not from a teacher model. Each task is62capped at 30k training examples, so no task dominates. Many tasks aren't in FLAN v2,63for example DynaSent, DynaHate, discriminative bAbI, epistemic logic, RuleTaker,64veridicality and dozens of NLI datasets.65 66```python67from datasets import load_dataset68 69ds = load_dataset("tasksource/tasksource-instruct", split="train")70ds = ds.filter(lambda use: use == "commercial", input_columns="license_use") # optional71```72 73## Format74 75| column | content |76|---|---|77| `inputs` | the instruction, the example, and the answer options |78| `targets` | the answer: an option (`entailment.`), a letter (`B.`), or `word: TAG` lines for token tasks |79| `task` | the tasksource task id |80| `license`, `license_use` | the source's licenses, see below |81 82Prompts ask for the answer with no explanation, so the short targets don't teach a83model to stop explaining in general. Tasks are interleaved round-robin, so any84slice of the split mixes them. Validation and test keep up to 500 examples per task.85 86`tasksource-instruct` works well mixed with FLAN v2 or other instruction data. It87covers discriminative reasoning tasks that those sets cover less.88 89For preference pairs built from the same rows, see90[tasksource_dpo_pairs](https://huggingface.co/datasets/tasksource/tasksource_dpo_pairs).91For soft labels, ratings and multi-question requests, see92[tasksource-jev-typed-decisions](https://huggingface.co/datasets/tasksource/tasksource-jev-typed-decisions).93 94## Reproducibility95 96The dataset is built by97[`scripts/build_instruct_dataset.py`](https://github.com/sileod/tasksource/blob/main/scripts/build_instruct_dataset.py):98 99```bash100PYTHONPATH=.:src python scripts/build_instruct_dataset.py --finalize101```102 103Sources are loaded at pinned Hub revisions. [sources.yaml](sources.yaml) records, per104task, the Hub dataset, revision, original dataset, licenses and row counts, and105`build-report.jsonl` records the code commit of each task's build. MMLU, BIG-bench and106BLiMP are left out, so they stay clean for evaluation. Other public benchmarks (GLUE,107SuperGLUE, HellaSwag, PIQA, ...) are **in** the data through their training splits.108 109## License and scope110 111Tasksource harmonizes datasets from many publishers; their original licenses112and terms still apply, hence `license: other`.113 114- `license` lists the `license` of the Hub dataset card the task was loaded from,115 and of the original dataset behind a tasksource copy. It also lists licenses recorded116 by the [Data Provenance Initiative](https://www.dataprovenance.org/), marked `(DPI)`.117- `license_use` takes the most restrictive of those: `non-commercial` if any is118 non-commercial or academic-only, `commercial` if one allows commercial use (share-alike119 and copyleft included), and `unspecified` otherwise.120 121This is a best-effort aid, not legal advice. Check the original terms before relying on them.122 123## Citation124 125```bibtex126@inproceedings{sileo-2024-tasksource,127 title = "tasksource: A Large Collection of {NLP} tasks with a Structured Dataset Preprocessing Framework",128 author = "Sileo, Damien",129 booktitle = "Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)",130 month = may,131 year = "2024",132 address = "Torino, Italia",133 publisher = "ELRA and ICCL",134 url = "https://aclanthology.org/2024.lrec-main.1361/",135 pages = "15655--15684",136}137```138 