Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Compactbot /slm-parameter-audit SLM card-vs-artifact parameter audit An autonomous audit of small-language-model repos on the Hugging Face Hub. For each in-scope model (independent builders training very small models from scratch, roughly 0.5M–500M parameters), the parameter count stated in the model card is compared against the actual artifact: the safetensors header, config.json, and the training script where present. A mismatch is recorded when the card's number does not match the artifact's real parameter… See the full description on the dataset page: https://huggingface.co/datasets/Compactbot/slm-parameter-audit.text-generation2 likes2.8k downloads2d agoHugging Face02prithivMLmods /Gargantua-R1-Compact Gargantua-R1 Distribution Gargantua-R1-Compact(experimental purpose) Gargantua-R1-Compact is a large-scale, high-quality reasoning dataset primarily designed for mathematical reasoning and STEM education. It contains approximately 6.67 million problems and solution traces, with a strong emphasis on mathematics (over 70%), as well as coverage of scientific domains, algorithmic challenges, and creative logic puzzles. The dataset is suitable for training and evaluating… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Gargantua-R1-Compact.texttext-generation1M<n<10M7 likes440 downloads1y agoHugging Face03moebiusT7 /l0-essentials-compact-ablation L0 Essentials compact ablation — raw rows Row-level outputs from ablating the MOBIUS MMV L0 Essentials governance prompt on six local models, three axes (tool loop / false premise / abstain chat), with the scorer sources and the pre-registered predictions. Status: experimental; not adversarially reviewed. The scorers had five documented defects during the work (all fixed, rows rescored); raw outputs are included so you can rescore with your own instrument. Headline… See the full description on the dataset page: https://huggingface.co/datasets/moebiusT7/l0-essentials-compact-ablation.text-generation0 likes112 downloads29d agoHugging Face04schneiderkamplab /dfm14-dala-v2-gl-compact DaLA v2 — Galician — DFM14 compact audit subset One language dataset combining all selected base and additive runs. Two task views share the same accepted source records; they are not independent observations. Each independently audited clean control contributes one correct input per task; each passing pair contributes one corrupted input per task. No extra clean controls are synthesized. Split/view Acceptability rows Correction rows train_representative 807,437 807… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm14-dala-v2-gl-compact.texttext-classification1M<n<10M0 likes65 downloads4d agoHugging Face05schneiderkamplab /dfm14-dala-v2-tr-compact DaLA v2 — Turkish — DFM14 compact audit subset One language dataset combining all selected base and additive runs. Two task views share the same accepted source records; they are not independent observations. Each independently audited clean control contributes one correct input per task; each passing pair contributes one corrupted input per task. No extra clean controls are synthesized. Split/view Acceptability rows Correction rows train_representative 920,794 920… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm14-dala-v2-tr-compact.texttext-classification1M<n<10M0 likes65 downloads4d agoHugging Face06schneiderkamplab /dfm14-dala-v2-ar-compact DaLA v2 — Modern Standard Arabic — DFM14 compact audit subset One language dataset combining all selected base and additive runs. Two task views share the same accepted source records; they are not independent observations. Each independently audited clean control contributes one correct input per task; each passing pair contributes one corrupted input per task. No extra clean controls are synthesized. Split/view Acceptability rows Correction rows train_representative… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm14-dala-v2-ar-compact.texttext-classification1M<n<10M0 likes65 downloads4d agoHugging Face07schneiderkamplab /dfm14-dala-v2-zh-compact DaLA v2 — Chinese — DFM14 compact audit subset One language dataset combining all selected base and additive runs. Two task views share the same accepted source records; they are not independent observations. Each independently audited clean control contributes one correct input per task; each passing pair contributes one corrupted input per task. No extra clean controls are synthesized. Split/view Acceptability rows Correction rows train_representative 1,037,788 1… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm14-dala-v2-zh-compact.texttext-classification1M<n<10M0 likes65 downloads4d agoHugging Face08introvoyz041 /Gargantua-R1-Compact Gargantua-R1 Distribution Gargantua-R1-Compact(experimental purpose) Gargantua-R1-Compact is a large-scale, high-quality reasoning dataset primarily designed for mathematical reasoning and STEM education. It contains approximately 6.67 million problems and solution traces, with a strong emphasis on mathematics (over 70%), as well as coverage of scientific domains, algorithmic challenges, and creative logic puzzles. The dataset is suitable for training and evaluating… See the full description on the dataset page: https://huggingface.co/datasets/introvoyz041/Gargantua-R1-Compact.texttext-generation1M<n<10M0 likes64 downloads8mo agoHugging Face09schneiderkamplab /dfm14-dala-v2-ga-compact DaLA v2 — Irish — DFM14 compact audit subset One language dataset combining all selected base and additive runs. Two task views share the same accepted source records; they are not independent observations. Each independently audited clean control contributes one correct input per task; each passing pair contributes one corrupted input per task. No extra clean controls are synthesized. Split/view Acceptability rows Correction rows train_representative 824,204 824,204… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm14-dala-v2-ga-compact.texttext-classification1M<n<10M0 likes64 downloads4d agoHugging Face10schneiderkamplab /dfm14-dala-v2-mk-compact DaLA v2 — Macedonian — DFM14 compact audit subset One language dataset combining all selected base and additive runs. Two task views share the same accepted source records; they are not independent observations. Each independently audited clean control contributes one correct input per task; each passing pair contributes one corrupted input per task. No extra clean controls are synthesized. Split/view Acceptability rows Correction rows train_representative 795,977 795… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm14-dala-v2-mk-compact.texttext-classification1M<n<10M0 likes63 downloads4d agoHugging Face11schneiderkamplab /dfm14-dala-v2-cy-compact DaLA v2 — Welsh — DFM14 compact audit subset One language dataset combining all selected base and additive runs. Two task views share the same accepted source records; they are not independent observations. Each independently audited clean control contributes one correct input per task; each passing pair contributes one corrupted input per task. No extra clean controls are synthesized. Split/view Acceptability rows Correction rows train_representative 1,046,231 1,046… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm14-dala-v2-cy-compact.texttext-classification1M<n<10M0 likes63 downloads4d agoHugging Face12schneiderkamplab /dfm14-dala-v2-vi-compact DaLA v2 — Vietnamese — DFM14 compact audit subset One language dataset combining all selected base and additive runs. Two task views share the same accepted source records; they are not independent observations. Each independently audited clean control contributes one correct input per task; each passing pair contributes one corrupted input per task. No extra clean controls are synthesized. Split/view Acceptability rows Correction rows train_representative 859,656 859… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm14-dala-v2-vi-compact.texttext-classification1M<n<10M0 likes63 downloads4d agoHugging Face13schneiderkamplab /dfm14-dala-v2-ru-compact DaLA v2 — Russian — DFM14 compact audit subset One language dataset combining all selected base and additive runs. Two task views share the same accepted source records; they are not independent observations. Each independently audited clean control contributes one correct input per task; each passing pair contributes one corrupted input per task. No extra clean controls are synthesized. Split/view Acceptability rows Correction rows train_representative 823,728 823… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm14-dala-v2-ru-compact.texttext-classification1M<n<10M0 likes62 downloads4d agoHugging Face14schneiderkamplab /dfm14-dala-v2-ja-compact DaLA v2 — Japanese — DFM14 compact audit subset One language dataset combining all selected base and additive runs. Two task views share the same accepted source records; they are not independent observations. Each independently audited clean control contributes one correct input per task; each passing pair contributes one corrupted input per task. No extra clean controls are synthesized. Split/view Acceptability rows Correction rows train_representative 911,381 911… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm14-dala-v2-ja-compact.texttext-classification1M<n<10M0 likes61 downloads4d agoHugging Face15schneiderkamplab /dfm14-dala-v2-he-compact DaLA v2 — Hebrew — DFM14 compact audit subset One language dataset combining all selected base and additive runs. Two task views share the same accepted source records; they are not independent observations. Each independently audited clean control contributes one correct input per task; each passing pair contributes one corrupted input per task. No extra clean controls are synthesized. Split/view Acceptability rows Correction rows train_representative 1,049,799 1,049… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm14-dala-v2-he-compact.texttext-classification1M<n<10M0 likes61 downloads4d agoHugging Face16schneiderkamplab /dfm14-dala-v2-eu-compact DaLA v2 — Basque — DFM14 compact audit subset One language dataset combining all selected base and additive runs. Two task views share the same accepted source records; they are not independent observations. Each independently audited clean control contributes one correct input per task; each passing pair contributes one corrupted input per task. No extra clean controls are synthesized. Split/view Acceptability rows Correction rows train_representative 887,025 887,025… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm14-dala-v2-eu-compact.texttext-classification1M<n<10M0 likes59 downloads4d agoHugging Face17schneiderkamplab /dfm14-dala-v2-mt-compact DaLA v2 — Maltese — DFM14 compact audit subset One language dataset combining all selected base and additive runs. Two task views share the same accepted source records; they are not independent observations. Each independently audited clean control contributes one correct input per task; each passing pair contributes one corrupted input per task. No extra clean controls are synthesized. Split/view Acceptability rows Correction rows train_representative 807,999 807… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm14-dala-v2-mt-compact.texttext-classification1M<n<10M0 likes59 downloads4d agoHugging Face18schneiderkamplab /dfm14-dala-v2-id-compact DaLA v2 — Indonesian — DFM14 compact audit subset One language dataset combining all selected base and additive runs. Two task views share the same accepted source records; they are not independent observations. Each independently audited clean control contributes one correct input per task; each passing pair contributes one corrupted input per task. No extra clean controls are synthesized. Split/view Acceptability rows Correction rows train_representative 879,440 879… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm14-dala-v2-id-compact.texttext-classification1M<n<10M0 likes58 downloads4d agoHugging Face19schneiderkamplab /dfm14-dala-v2-ko-compact DaLA v2 — Korean — DFM14 compact audit subset One language dataset combining all selected base and additive runs. Two task views share the same accepted source records; they are not independent observations. Each independently audited clean control contributes one correct input per task; each passing pair contributes one corrupted input per task. No extra clean controls are synthesized. Split/view Acceptability rows Correction rows train_representative 833,536 833,536… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm14-dala-v2-ko-compact.texttext-classification1M<n<10M0 likes58 downloads4d agoHugging Face20schneiderkamplab /dfm14-dala-v2-hi-compact DaLA v2 — Hindi — DFM14 compact audit subset One language dataset combining all selected base and additive runs. Two task views share the same accepted source records; they are not independent observations. Each independently audited clean control contributes one correct input per task; each passing pair contributes one corrupted input per task. No extra clean controls are synthesized. Split/view Acceptability rows Correction rows train_representative 1,038,773 1,038… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm14-dala-v2-hi-compact.text-classification0 likes54 downloads4d agoHugging Face21schneiderkamplab /dfm13-dala-v2-lv-compact dfm13-dala-v2-lv-compact One language dataset combining accepted baseline and recovery pools where available. Exact native conversations are deduplicated within each task/split; split conflicts fail closed. Full native messages, target indices and source provenance are preserved; publication_pool records origin. Task configs separate acceptability and correction; heldout views never become train. Automated compact audits are not native-speaker or producer per-edit certification.… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm13-dala-v2-lv-compact.texttext-generation1M<n<10M0 likes49 downloads6d agoHugging Face22schneiderkamplab /dfm13-dala-v2-pt-PT-compact dfm13-dala-v2-pt-PT-compact One language dataset combining accepted baseline and recovery pools where available. Exact native conversations are deduplicated within each task/split; split conflicts fail closed. Full native messages, target indices and source provenance are preserved; publication_pool records origin. Task configs separate acceptability and correction; heldout views never become train. Automated compact audits are not native-speaker or producer per-edit… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm13-dala-v2-pt-PT-compact.texttext-generation1M<n<10M0 likes48 downloads6d agoHugging Face23schneiderkamplab /dfm13-dala-v2-en-compact dfm13-dala-v2-en-compact One language dataset combining accepted baseline and recovery pools where available. Exact native conversations are deduplicated within each task/split; split conflicts fail closed. Full native messages, target indices and source provenance are preserved; publication_pool records origin. Task configs separate acceptability and correction; heldout views never become train. Automated compact audits are not native-speaker or producer per-edit certification.… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm13-dala-v2-en-compact.texttext-generation1M<n<10M0 likes47 downloads6d agoHugging Face24schneiderkamplab /dfm13-dala-v2-lb-compact dfm13-dala-v2-lb-compact One language dataset combining accepted baseline and recovery pools where available. Exact native conversations are deduplicated within each task/split; split conflicts fail closed. Full native messages, target indices and source provenance are preserved; publication_pool records origin. Task configs separate acceptability and correction; heldout views never become train. Automated compact audits are not native-speaker or producer per-edit certification.… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm13-dala-v2-lb-compact.texttext-generation10K<n<100K0 likes47 downloads6d agoHugging Face25schneiderkamplab /dfm13-dala-v2-lt-compact dfm13-dala-v2-lt-compact One language dataset combining accepted baseline and recovery pools where available. Exact native conversations are deduplicated within each task/split; split conflicts fail closed. Full native messages, target indices and source provenance are preserved; publication_pool records origin. Task configs separate acceptability and correction; heldout views never become train. Automated compact audits are not native-speaker or producer per-edit certification.… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm13-dala-v2-lt-compact.texttext-generation1M<n<10M0 likes46 downloads6d agoHugging Face26schneiderkamplab /dfm13-dala-v2-bs-compact dfm13-dala-v2-bs-compact One language dataset combining accepted baseline and recovery pools where available. Exact native conversations are deduplicated within each task/split; split conflicts fail closed. Full native messages, target indices and source provenance are preserved; publication_pool records origin. Task configs separate acceptability and correction; heldout views never become train. Automated compact audits are not native-speaker or producer per-edit certification.… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm13-dala-v2-bs-compact.texttext-generation100K<n<1M0 likes44 downloads6d agoHugging Face27harvardairobotics /MedConclusion-Compact MedConclusion-Compact MedConclusion is a large-scale dataset of 5.7M PubMed structured abstracts for biomedical conclusion generation. Each instance pairs the non-conclusion sections of an abstract with the original author-written conclusion, providing naturally occurring supervision for evidence-to-conclusion reasoning. MedConclusion also includes journal-level metadata such as biomedical category and SJR, enabling subgroup analysis across biomedical domains. This repository… See the full description on the dataset page: https://huggingface.co/datasets/harvardairobotics/MedConclusion-Compact.tabulartext-generation100K<n<1M2 likes43 downloads6mo agoHugging Face28schneiderkamplab /dfm13-dala-v2-sr-compact dfm13-dala-v2-sr-compact One language dataset combining accepted baseline and recovery pools where available. Exact native conversations are deduplicated within each task/split; split conflicts fail closed. Full native messages, target indices and source provenance are preserved; publication_pool records origin. Task configs separate acceptability and correction; heldout views never become train. Automated compact audits are not native-speaker or producer per-edit certification.… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm13-dala-v2-sr-compact.text-generation0 likes41 downloads6d agoHugging Face29schneiderkamplab /dfm13-dala-v2-fa-compact dfm13-dala-v2-fa-compact One language dataset combining accepted baseline and recovery pools where available. Exact native conversations are deduplicated within each task/split; split conflicts fail closed. Full native messages, target indices and source provenance are preserved; publication_pool records origin. Task configs separate acceptability and correction; heldout views never become train. Automated compact audits are not native-speaker or producer per-edit certification.… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm13-dala-v2-fa-compact.text-generation0 likes40 downloads6d agoHugging Face30schneiderkamplab /dfm13-dala-v2-pl-compact dfm13-dala-v2-pl-compact One language dataset combining accepted baseline and recovery pools where available. Exact native conversations are deduplicated within each task/split; split conflicts fail closed. Full native messages, target indices and source provenance are preserved; publication_pool records origin. Task configs separate acceptability and correction; heldout views never become train. Automated compact audits are not native-speaker or producer per-edit certification.… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm13-dala-v2-pl-compact.text-generation0 likes39 downloads6d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.