Lelonthecodeur/multi-task-dataset
Multi-Task Dataset Description A large-scale multi-task dataset designed for training and evaluating AI models across reasoning, mathematics, code, research, verification, data analysis, and general problem solving. Content 100,000,001 examples 20+ task families English + French Train / Validation / Test splits Structured reasoning and verification signals Multiple difficulty levels OOD and generalization-oriented examples Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Lelonthecodeur/multi-task-dataset.
2228
1---2license: mit3tags:4- code5- Text6- Science7- Math8- Logic9---10 11# Multi-Task Dataset12 13## Description14 15A large-scale multi-task dataset designed for training and evaluating AI models across **reasoning, mathematics, code, research, verification, data analysis, and general problem solving**.16 17## Content18 19* **100,000,001** examples20* **20+ task families**21* English + French22* Train / Validation / Test splits23* Structured reasoning and verification signals24* Multiple difficulty levels25* OOD and generalization-oriented examples26 27## Dataset Structure28 29| Split | Percentage | Examples |30| ---------- | ---------: | --------------: |31| Train | 97.999999% | 98,000,000 |32| Validation | 0.999999% | 1,000,000 |33| Test | 1.000000% | 1,000,001 |34| **Total** | **100%** | **100,000,001** |35 36## Task Categories37 38| Category | Share | Examples |39| ------------------------ | ----: | ---------: |40| Mathematics | 10% | 10,000,000 |41| Logic | 8% | 8,000,000 |42| Programming | 10% | 10,000,000 |43| Data Analysis | 8% | 8,000,000 |44| Reasoning | 10% | 10,000,000 |45| Research | 6% | 6,000,000 |46| Fact Checking | 5% | 5,000,000 |47| Self-Correction | 6% | 6,000,000 |48| Instruction Following | 6% | 6,000,000 |49| Planning | 5% | 5,000,000 |50| Constraint Reasoning | 4% | 4,000,000 |51| Counterexample Reasoning | 4% | 4,000,000 |52| Adversarial Reasoning | 4% | 4,000,000 |53| Calibration | 3% | 3,000,000 |54| Ambiguity Handling | 3% | 3,000,000 |55| Error Analysis | 3% | 3,000,000 |56| Generalization | 3% | 3,000,000 |57| Prompt Review | 2% | 2,000,000 |58| Consistency Checking | 2% | 2,000,000 |59| Evidence Checking | 1% | 1,000,000 |60 61## Main Capabilities62 63The dataset is designed to improve:64 65* Mathematical reasoning66* Logical reasoning67* Code generation68* Code understanding69* Data analysis70* Research methodology71* Error detection72* Error correction73* Self-checking74* Instruction following75* Constraint satisfaction76* Counterexample detection77* Ambiguity resolution78* Prompt consistency79* Confidence estimation80* Uncertainty handling81* Generalization82* Verification83 84## Difficulty85 86Examples are distributed across multiple difficulty levels:87 88* `easy`89* `medium`90* `hard`91* `very_hard`92* `extreme`93 94## Verification Signals95 96Examples can contain structured fields for:97 98* `math_check`99* `logic_check`100* `code_check`101* `data_analysis_check`102* `constraint_check`103* `consistency_check`104* `counterexample_check`105* `evidence_check`106* `source_check`107* `error_detection`108* `error_repair`109* `prompt_review`110* `prompt_alignment`111* `confidence`112* `uncertainty`113* `answerability`114 115## Usage116 117Install the required library:118 119```bash120pip install -U datasets121```122 123Load the dataset:124 125```python126from datasets import load_dataset127 128dataset = load_dataset(129 "Lelonthecodeur/multi-task-dataset",130 streaming=True131)132 133train = dataset["train"]134 135for example in train:136 print(example)137 break138```139 140Load a specific split:141 142```python143from datasets import load_dataset144 145train = load_dataset(146 "Lelonthecodeur/multi-task-dataset",147 split="train",148 streaming=True149)150```151 152Streaming is recommended for the full dataset because of its size.153 154## Hugging Face CLI155 156Login:157 158```bash159hf auth login160```161 162Clone:163 164```bash165git lfs install166git clone https://huggingface.co/datasets/Lelonthecodeur/multi-task-dataset167```168 169Push an update:170 171```bash172cd multi-task-dataset173git add .174git commit -m "Update dataset"175git push176```177 178## Python Upload179 180```python181from huggingface_hub import HfApi182 183api = HfApi(token="YOUR_HF_TOKEN")184 185api.upload_folder(186 folder_path="/kaggle/working/multi-task-dataset",187 repo_id="Lelonthecodeur/multi-task-dataset",188 repo_type="dataset",189 commit_message="Update dataset",190)191```192 193## Data Format194 195The dataset is stored in **Parquet** format.196 197Main fields include:198 199```text200id201task_family202task_type203domain204difficulty205language206instruction207context208response209analysis_plan210verification211prompt_review212prompt_alignment213constraint_check214consistency_check215math_check216logic_check217counterexample_check218data_analysis_check219code_check220evidence_check221source_check222hallucination_control223error_detection224error_repair225answerability226confidence227uncertainty228reasoning_depth229minimal_sufficient_reasoning230unnecessary_reasoning231stop_condition232surface_variation233numeric_variation234structure_variation235ood_style236quality_score237generator_version238```239 240## Future Updates241 242### V2 — Robustness243 244Planned improvements:245 246* Harder reasoning tasks247* Adversarial examples248* Hard negatives249* Better deduplication250* Near-duplicate detection251* Leakage detection252* Stronger OOD splits253* Better generalization testing254 255### V3 — Science & Research256 257Planned additions:258 259* Scientific reasoning260* Scientific knowledge261* Research methodology262* Experimental design263* Hypothesis evaluation264* Scientific data analysis265* Evidence comparison266* Source comparison267* Uncertainty analysis268 269### V4 — Mega Deep270 271Planned addition of approximately **10M highly difficult examples**.272 273The objective is to target specific weaknesses found during model evaluation instead of simply increasing prompt complexity.274 275```text276Model277 ↓278Benchmark279 ↓280Failure Detection281 ↓282Weak Skill Detection283 ↓284Targeted Hard Examples285 ↓286Verification287 ↓288Deduplication289 ↓290OOD / Adversarial Tests291 ↓292Training293 ↓294New Benchmark295```296 297### V5 — Science × Knowledge × Logic × Experience298 299Future expansion combining:300 301* Science302* Knowledge303* Complex logic304* Experience-based problem solving305* Cross-domain reasoning306* Multi-step verification307* Novel situations308* Adaptive evaluation309 310## Font311 312For standard text:313 314```python315import matplotlib.pyplot as plt316 317plt.rcParams["font.family"] = "DejaVu Sans"318```319 320For multilingual text:321 322```python323import matplotlib.pyplot as plt324 325plt.rcParams["font.family"] = ["Noto Sans", "Noto Sans CJK JP"]326```327 328## License329 330MIT License331 332Copyright (c) 2026 Lelonthecodeur333 334Permission is hereby granted, free of charge, to any person obtaining a copy of this dataset and associated files, to use, copy, modify, merge, publish, distribute, sublicense, and sell copies of the dataset, subject to the conditions of the MIT License.335 336## Version337 338**Current version:** `v1.1`339 340**Total examples:** `100,000,001`341 342**Format:** Parquet343 344**Status:** Active development345 346## Citation347 348```bibtex349@dataset{multi_task_dataset,350 title = {multi-task-dataset},351 author = {Lelonthecodeur},352 year = {2026},353 publisher = {Hugging Face},354 version = {1.1},355 note = {100,000,001 examples}356}357```