Team Ai
Datasetpublic

Lelonthecodeur/multi-task-dataset

Multi-Task Dataset Description A large-scale multi-task dataset designed for training and evaluating AI models across reasoning, mathematics, code, research, verification, data analysis, and general problem solving. Content 100,000,001 examples 20+ task families English + French Train / Validation / Test splits Structured reasoning and verification signals Multiple difficulty levels OOD and generalization-oriented examples Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Lelonthecodeur/multi-task-dataset.

sourceHugging Facemitupdated 23d agoView on Hugging Face
2likes228downloads
README.md357 linesDownload Raw Back to root
1---2license: mit3tags:4- code5- Text6- Science7- Math8- Logic9---10 11# Multi-Task Dataset12 13## Description14 15A large-scale multi-task dataset designed for training and evaluating AI models across **reasoning, mathematics, code, research, verification, data analysis, and general problem solving**.16 17## Content18 19* **100,000,001** examples20* **20+ task families**21* English + French22* Train / Validation / Test splits23* Structured reasoning and verification signals24* Multiple difficulty levels25* OOD and generalization-oriented examples26 27## Dataset Structure28 29| Split      | Percentage |        Examples |30| ---------- | ---------: | --------------: |31| Train      | 97.999999% |      98,000,000 |32| Validation |  0.999999% |       1,000,000 |33| Test       |  1.000000% |       1,000,001 |34| **Total**  |   **100%** | **100,000,001** |35 36## Task Categories37 38| Category                 | Share |   Examples |39| ------------------------ | ----: | ---------: |40| Mathematics              |   10% | 10,000,000 |41| Logic                    |    8% |  8,000,000 |42| Programming              |   10% | 10,000,000 |43| Data Analysis            |    8% |  8,000,000 |44| Reasoning                |   10% | 10,000,000 |45| Research                 |    6% |  6,000,000 |46| Fact Checking            |    5% |  5,000,000 |47| Self-Correction          |    6% |  6,000,000 |48| Instruction Following    |    6% |  6,000,000 |49| Planning                 |    5% |  5,000,000 |50| Constraint Reasoning     |    4% |  4,000,000 |51| Counterexample Reasoning |    4% |  4,000,000 |52| Adversarial Reasoning    |    4% |  4,000,000 |53| Calibration              |    3% |  3,000,000 |54| Ambiguity Handling       |    3% |  3,000,000 |55| Error Analysis           |    3% |  3,000,000 |56| Generalization           |    3% |  3,000,000 |57| Prompt Review            |    2% |  2,000,000 |58| Consistency Checking     |    2% |  2,000,000 |59| Evidence Checking        |    1% |  1,000,000 |60 61## Main Capabilities62 63The dataset is designed to improve:64 65* Mathematical reasoning66* Logical reasoning67* Code generation68* Code understanding69* Data analysis70* Research methodology71* Error detection72* Error correction73* Self-checking74* Instruction following75* Constraint satisfaction76* Counterexample detection77* Ambiguity resolution78* Prompt consistency79* Confidence estimation80* Uncertainty handling81* Generalization82* Verification83 84## Difficulty85 86Examples are distributed across multiple difficulty levels:87 88* `easy`89* `medium`90* `hard`91* `very_hard`92* `extreme`93 94## Verification Signals95 96Examples can contain structured fields for:97 98* `math_check`99* `logic_check`100* `code_check`101* `data_analysis_check`102* `constraint_check`103* `consistency_check`104* `counterexample_check`105* `evidence_check`106* `source_check`107* `error_detection`108* `error_repair`109* `prompt_review`110* `prompt_alignment`111* `confidence`112* `uncertainty`113* `answerability`114 115## Usage116 117Install the required library:118 119```bash120pip install -U datasets121```122 123Load the dataset:124 125```python126from datasets import load_dataset127 128dataset = load_dataset(129    "Lelonthecodeur/multi-task-dataset",130    streaming=True131)132 133train = dataset["train"]134 135for example in train:136    print(example)137    break138```139 140Load a specific split:141 142```python143from datasets import load_dataset144 145train = load_dataset(146    "Lelonthecodeur/multi-task-dataset",147    split="train",148    streaming=True149)150```151 152Streaming is recommended for the full dataset because of its size.153 154## Hugging Face CLI155 156Login:157 158```bash159hf auth login160```161 162Clone:163 164```bash165git lfs install166git clone https://huggingface.co/datasets/Lelonthecodeur/multi-task-dataset167```168 169Push an update:170 171```bash172cd multi-task-dataset173git add .174git commit -m "Update dataset"175git push176```177 178## Python Upload179 180```python181from huggingface_hub import HfApi182 183api = HfApi(token="YOUR_HF_TOKEN")184 185api.upload_folder(186    folder_path="/kaggle/working/multi-task-dataset",187    repo_id="Lelonthecodeur/multi-task-dataset",188    repo_type="dataset",189    commit_message="Update dataset",190)191```192 193## Data Format194 195The dataset is stored in **Parquet** format.196 197Main fields include:198 199```text200id201task_family202task_type203domain204difficulty205language206instruction207context208response209analysis_plan210verification211prompt_review212prompt_alignment213constraint_check214consistency_check215math_check216logic_check217counterexample_check218data_analysis_check219code_check220evidence_check221source_check222hallucination_control223error_detection224error_repair225answerability226confidence227uncertainty228reasoning_depth229minimal_sufficient_reasoning230unnecessary_reasoning231stop_condition232surface_variation233numeric_variation234structure_variation235ood_style236quality_score237generator_version238```239 240## Future Updates241 242### V2 — Robustness243 244Planned improvements:245 246* Harder reasoning tasks247* Adversarial examples248* Hard negatives249* Better deduplication250* Near-duplicate detection251* Leakage detection252* Stronger OOD splits253* Better generalization testing254 255### V3 — Science & Research256 257Planned additions:258 259* Scientific reasoning260* Scientific knowledge261* Research methodology262* Experimental design263* Hypothesis evaluation264* Scientific data analysis265* Evidence comparison266* Source comparison267* Uncertainty analysis268 269### V4 — Mega Deep270 271Planned addition of approximately **10M highly difficult examples**.272 273The objective is to target specific weaknesses found during model evaluation instead of simply increasing prompt complexity.274 275```text276Model277  ↓278Benchmark279  ↓280Failure Detection281  ↓282Weak Skill Detection283  ↓284Targeted Hard Examples285  ↓286Verification287  ↓288Deduplication289  ↓290OOD / Adversarial Tests291  ↓292Training293  ↓294New Benchmark295```296 297### V5 — Science × Knowledge × Logic × Experience298 299Future expansion combining:300 301* Science302* Knowledge303* Complex logic304* Experience-based problem solving305* Cross-domain reasoning306* Multi-step verification307* Novel situations308* Adaptive evaluation309 310## Font311 312For standard text:313 314```python315import matplotlib.pyplot as plt316 317plt.rcParams["font.family"] = "DejaVu Sans"318```319 320For multilingual text:321 322```python323import matplotlib.pyplot as plt324 325plt.rcParams["font.family"] = ["Noto Sans", "Noto Sans CJK JP"]326```327 328## License329 330MIT License331 332Copyright (c) 2026 Lelonthecodeur333 334Permission is hereby granted, free of charge, to any person obtaining a copy of this dataset and associated files, to use, copy, modify, merge, publish, distribute, sublicense, and sell copies of the dataset, subject to the conditions of the MIT License.335 336## Version337 338**Current version:** `v1.1`339 340**Total examples:** `100,000,001`341 342**Format:** Parquet343 344**Status:** Active development345 346## Citation347 348```bibtex349@dataset{multi_task_dataset,350  title        = {multi-task-dataset},351  author       = {Lelonthecodeur},352  year         = {2026},353  publisher    = {Hugging Face},354  version      = {1.1},355  note         = {100,000,001 examples}356}357```