Team Ai
Datasetpublic

tasksource/tasksource-instruct

tasksource-instruct Instruction-tuning data recast from the ~480 English classification, multiple-choice and token-classification tasks of tasksource. Every example comes from a human-built dataset (NLI, logical reasoning, sentiment, hate speech, discourse, argumentation, ...), not from a teacher model. Each task is capped at 30k training examples, so no task dominates. Many tasks aren't in FLAN v2, for example DynaSent, DynaHate, discriminative bAbI, epistemic logic, RuleTaker… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/tasksource-instruct.

sourceHugging Faceotherupdated 13d agoView on Hugging Face
24likes2.6kdownloads
README.md138 linesDownload Raw Back to root
1---2pretty_name: tasksource-instruct3language:4- en5license: other6size_categories:7- 1M<n<10M8task_categories:9- text-generation10- text-classification11- token-classification12- zero-shot-classification13tags:14- instructions15- instruction-tuning16- instruction-finetuning17- flan18- promptsource19- tasksource20configs:21- config_name: default22  data_files:23  - split: train24    path: data/train-*25  - split: test26    path: data/test-*27  - split: validation28    path: data/validation-*29dataset_info:30  features:31  - name: inputs32    dtype: string33  - name: targets34    dtype: string35  - name: task36    dtype: string37  - name: license38    dtype: string39  - name: license_use40    dtype: string41  splits:42  - name: train43    num_bytes: 650981849144    num_examples: 728225545  - name: validation46    num_bytes: 18388851547    num_examples: 18491448  - name: test49    num_bytes: 18411635250    num_examples: 19283051  download_size: 312279200752  dataset_size: 687782335853---54 55# tasksource-instruct56 57**Instruction-tuning data recast from the ~480 English classification, multiple-choice58and token-classification tasks of [tasksource](https://github.com/sileod/tasksource).**59 60Every example comes from a human-built dataset (NLI, logical reasoning, sentiment,61hate speech, discourse, argumentation, ...), not from a teacher model. Each task is62capped at 30k training examples, so no task dominates. Many tasks aren't in FLAN v2,63for example DynaSent, DynaHate, discriminative bAbI, epistemic logic, RuleTaker,64veridicality and dozens of NLI datasets.65 66```python67from datasets import load_dataset68 69ds = load_dataset("tasksource/tasksource-instruct", split="train")70ds = ds.filter(lambda use: use == "commercial", input_columns="license_use")  # optional71```72 73## Format74 75| column | content |76|---|---|77| `inputs` | the instruction, the example, and the answer options |78| `targets` | the answer: an option (`entailment.`), a letter (`B.`), or `word: TAG` lines for token tasks |79| `task` | the tasksource task id |80| `license`, `license_use` | the source's licenses, see below |81 82Prompts ask for the answer with no explanation, so the short targets don't teach a83model to stop explaining in general. Tasks are interleaved round-robin, so any84slice of the split mixes them. Validation and test keep up to 500 examples per task.85 86`tasksource-instruct` works well mixed with FLAN v2 or other instruction data. It87covers discriminative reasoning tasks that those sets cover less.88 89For preference pairs built from the same rows, see90[tasksource_dpo_pairs](https://huggingface.co/datasets/tasksource/tasksource_dpo_pairs).91For soft labels, ratings and multi-question requests, see92[tasksource-jev-typed-decisions](https://huggingface.co/datasets/tasksource/tasksource-jev-typed-decisions).93 94## Reproducibility95 96The dataset is built by97[`scripts/build_instruct_dataset.py`](https://github.com/sileod/tasksource/blob/main/scripts/build_instruct_dataset.py):98 99```bash100PYTHONPATH=.:src python scripts/build_instruct_dataset.py --finalize101```102 103Sources are loaded at pinned Hub revisions. [sources.yaml](sources.yaml) records, per104task, the Hub dataset, revision, original dataset, licenses and row counts, and105`build-report.jsonl` records the code commit of each task's build. MMLU, BIG-bench and106BLiMP are left out, so they stay clean for evaluation. Other public benchmarks (GLUE,107SuperGLUE, HellaSwag, PIQA, ...) are **in** the data through their training splits.108 109## License and scope110 111Tasksource harmonizes datasets from many publishers; their original licenses112and terms still apply, hence `license: other`.113 114- `license` lists the `license` of the Hub dataset card the task was loaded from,115  and of the original dataset behind a tasksource copy. It also lists licenses recorded116  by the [Data Provenance Initiative](https://www.dataprovenance.org/), marked `(DPI)`.117- `license_use` takes the most restrictive of those: `non-commercial` if any is118  non-commercial or academic-only, `commercial` if one allows commercial use (share-alike119  and copyleft included), and `unspecified` otherwise.120 121This is a best-effort aid, not legal advice. Check the original terms before relying on them.122 123## Citation124 125```bibtex126@inproceedings{sileo-2024-tasksource,127    title = "tasksource: A Large Collection of {NLP} tasks with a Structured Dataset Preprocessing Framework",128    author = "Sileo, Damien",129    booktitle = "Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)",130    month = may,131    year = "2024",132    address = "Torino, Italia",133    publisher = "ELRA and ICCL",134    url = "https://aclanthology.org/2024.lrec-main.1361/",135    pages = "15655--15684",136}137```138