Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01TigreGotico /arabic-stem-lexicon Arabic Diacritized-Stem Lexicon An undiacritized Arabic surface form → its most frequent diacritized stem. Standard Arabic writes no short vowels, so anything that has to pronounce Arabic must first put them back. A neural diacritizer does that well on rare words, where inference is the only thing there is. On common words it is the wrong tool: which vowels كتاب carries is not a thing to be inferred, it is a thing to be looked up — and models get exactly these wrong, reading… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/arabic-stem-lexicon.tabulartext-to-speech100K<n<1M0 likes3.8k downloads3mo agoHugging Face02qingyangzhang /Natural-Reasoning-STEM-25Ktabular10K<n<100K0 likes697 downloads1y agoHugging Face03gary23ai /STEM2Crystal-Bench STEM2Crystal-Bench STEM2Crystal-Bench is the benchmark for the paper "From Noisy STEM to Crystal Structure: Evidence-Structure CoDiffusion under Composition Constraints" (Chen & You, KDD 2026, Oral), which introduces STEM2Crystal CoDiffusion (SCCD). It evaluates methods that reconstruct a crystal structure from a noisy STEM image when the composition is known. The release has a large synthetic set with controlled noise and a small set of real STEM images, with ground-truth CIFs… See the full description on the dataset page: https://huggingface.co/datasets/gary23ai/STEM2Crystal-Bench.imageimage-to-text1K<n<10K1 likes414 downloads4mo agoHugging Face04lfaviate /China-K12-STEM-10K-CoT-Reasoning K12-STEM-CoT-Chinese 1.54M Chinese K12 STEM problems with chain-of-thought solutions, 48% with diagrams. The largest structured Chinese math/physics/chemistry reasoning dataset. This is a curated sample (10,000 problems) of the full 1.54M dataset available via API. Full Dataset Access Access the full 1,540,000+ problems via API → This Sample Full API Total problems 10,025 1,540,000+ With CoT solutions 10,025 1,490,000+ With diagrams 6,093 740,000+… See the full description on the dataset page: https://huggingface.co/datasets/lfaviate/China-K12-STEM-10K-CoT-Reasoning.tabularquestion-answering10K<n<100K3 likes353 downloads7mo agoHugging Face05aeyxen /stem-diagrams STEM Diagrams 30,325 technical diagrams (block diagrams, schematics, flowcharts, architectures) extracted from arXiv papers across six engineering fields, each with a source attribution and a quality score. Built by an LLM-curated pipeline and used to show that a small frozen-feature classifier can replace the paid LLM labeling gate. Paper: Distilling an LLM Diagram-Curation Pipeline into Local Classifiers (Adnan Abbasi, Thothica, 2026) Code:… See the full description on the dataset page: https://huggingface.co/datasets/aeyxen/stem-diagrams.imageimage-classification10K<n<100K0 likes284 downloads3mo agoHugging Face06lihaoxin2020 /ki-qwen-instruct-synthetic_1_stem_only-sft-temp0.6-on-mmlu_pro-0shot_cot-scillm-da553bdec9tabular1K<n<10K0 likes268 downloads7mo agoHugging Face07hanzla /STEM_Reasoningtabular10K<n<100K1 likes92 downloads2y agoHugging Face08a13905873166 /China-K12-STEM-10K-CoT-Reasoning K12-STEM-CoT-Chinese 1.54M Chinese K12 STEM problems with chain-of-thought solutions, 48% with diagrams. The largest structured Chinese math/physics/chemistry reasoning dataset. This is a curated sample (10,000 problems) of the full 1.54M dataset available via API. Full Dataset Access Access the full 1,540,000+ problems via API → This Sample Full API Total problems 10,025 1,540,000+ With CoT solutions 10,025 1,490,000+ With diagrams 6,093 740… See the full description on the dataset page: https://huggingface.co/datasets/a13905873166/China-K12-STEM-10K-CoT-Reasoning.tabularquestion-answering10K<n<100K1 likes83 downloads25d agoHugging Face09qingyangzhang /Natural-Reasoning-STEM-50Ktabular10K<n<100K0 likes75 downloads1y agoHugging Face10stemauro /multimodal-lucas Dataset card for Multi-modal LUCAS Dataset summary Multi-modal LUCAS aims at being a curated vision-language dataset from LUCAS survey data and in-situ field photos. LUCAS (Land Use/Cover Area Frame statistical Survey) is a land-monitoring exercise conducted by EUROSTAT in close cooperation with the Directorate-General responsible for Agriculture, with technical support from the Joint Research Centre (JRC). The survey has been repeated every three years since 2006… See the full description on the dataset page: https://huggingface.co/datasets/stemauro/multimodal-lucas.image100K<n<1M0 likes71 downloads1mo agoHugging Face11open-llm-leaderboard /Josephgflowers__Tinyllama-STEM-Cinder-Agent-v1-detailsgated Dataset Card for Evaluation run of Josephgflowers/Tinyllama-STEM-Cinder-Agent-v1 Dataset automatically created during the evaluation run of model Josephgflowers/Tinyllama-STEM-Cinder-Agent-v1 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Josephgflowers__Tinyllama-STEM-Cinder-Agent-v1-details.tabular10K<n<100K0 likes59 downloads2y agoHugging Face12StemSplitio /stem-separation-benchmark-2026 StemSplit Stem-Separation Benchmark 2026 A reproducible head-to-head comparison of every popular open-source music source-separation model against the StemSplit production API, evaluated on the standard MUSDB18-HQ test split using BSS Eval v4 and a small set of CC-BY tracks for qualitative listening. Built and maintained by the StemSplit team. Source code: scripts/hf-benchmark on GitHub. Leaderboard (median SDR per stem) model_id bass drums other vocals… See the full description on the dataset page: https://huggingface.co/datasets/StemSplitio/stem-separation-benchmark-2026.tabularaudio-to-audion<1K1 likes52 downloads5mo agoHugging Face13liujin99 /quadmix-stem-v1 QuaDMix-STEM v1: STEM-Focused Proxy Validation Set Script: scripts/validation_set/prepare_stem_v1.py HuggingFace: liujin99/quadmix-stem-v1 Files: stem_v1_tokenized.pt, stem_v1.parquet Overview STEM v1 is a validation set designed to focus the proxy model's optimization signal on STEM capabilities — mathematics, science knowledge, and logical reasoning. Unlike CAP v1 (broad capability coverage) or core_bmk (benchmark test format), STEM v1 uses only tasks that… See the full description on the dataset page: https://huggingface.co/datasets/liujin99/quadmix-stem-v1.tabularquestion-answering10K<n<100K0 likes52 downloads3mo agoHugging Face14Shann5 /bazi-hidden-stemsfrom datasets import load_dataset rows = load_dataset("Shann5/bazi-hidden-stems") # one row per (source, branch, stem) compare = load_dataset("Shann5/bazi-hidden-stems", "compare") # one row per branch, one column per source Mirror of hidden-stems/ in Shann5/bazi-open-data — the GitHub copy is canonical (JSON version, source texts, build). Corrections belong there as issues. Hidden stems (藏干) across sources Every earthly branch stores one to three heavenly… See the full description on the dataset page: https://huggingface.co/datasets/Shann5/bazi-hidden-stems.tabularn<1K0 likes47 downloads14d agoHugging Face15yen-av /tunix-stem-sft Reasoning Training Dataset for Tunix Competition Reasoning dataset for training 1-2B thinking models on math, coding, and science problems. Sources GSM8K: Grade school math with human reasoning traces TextbookReasoning: STEM problems with step-by-step solutions MBPP: Basic Python Programming prompts, with reasoning traces generated by gpt-oss-20b Format Each example contains: prompt: The problem statement reasoning: Step-by-step reasoning answer: Final… See the full description on the dataset page: https://huggingface.co/datasets/yen-av/tunix-stem-sft.tabular1M<n<10M0 likes42 downloads11mo agoHugging Face16tommymarto /STEM-wikipedia-22-12-en-embeddings-all-MiniLM-L6-v2tabular1M<n<10M1 likes40 downloads2y agoHugging Face17ToneCubeMedia /Pop-Rock-Hybrid-Stem-Dataset-cat001 Dataset Overview: Pop Rock Hybrid Stem Dataset (cat001) This dataset contains a curated collection of original instrumental music designed for commercial and research applications in music analysis, audio modeling, and production workflows. Every composition, arrangement, performance, sound design element, and production decision was created entirely through human musical and technical processes. All music contained in this dataset is 100% human-made (is_human_created: TRUE).… See the full description on the dataset page: https://huggingface.co/datasets/ToneCubeMedia/Pop-Rock-Hybrid-Stem-Dataset-cat001.tabularaudio-classificationn<1K0 likes40 downloads2mo agoHugging Face18Siesher /mits-stem-training-datatabular10K<n<100K0 likes39 downloads8mo agoHugging Face19liujin99 /quadmix-stem-v2 QuaDMix-STEM v2: STEM-Focused Proxy Validation Set with GPQA & MATH Script: scripts/validation_set/prepare_stem_v2.py HuggingFace: liujin99/quadmix-stem-v2 Files: stem_v2_tokenized.pt, stem_v2.parquet Overview STEM v2 is an upgraded validation set that fixes the two critical coverage gaps in STEM v1. In the v1 experiment, QuaDMix lost to Random downstream (CORE 0.1530 vs 0.1615), and root-cause analysis revealed: gpqa_diamond had no direct proxy — mapped from… See the full description on the dataset page: https://huggingface.co/datasets/liujin99/quadmix-stem-v2.tabularquestion-answering10K<n<100K0 likes38 downloads3mo agoHugging Face20electricsheepafrica /africa-ilo-emp-stem-sex-how-nb-employment-in-stem-occupations-by-sex-and-weekly-h Employment in STEM occupations by sex and weekly hours actually worked (thousands) | Africa (ILOSTAT) | Africa (Electric Sheep Africa metadata inventory) Size category: 1K<n<10K - Formats: parquet - Sector: economics_finance - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-ilo-emp-stem-sex-how-nb-employment-in-stem-occupations-by-sex-and-weekly-h.tabulartabular-classification1K<n<10K0 likes36 downloads2mo agoHugging Face21InfoBayAI /Indonesian-STEM-Textbook-DatasetgatedDataset Description: This dataset is a large-scale collection of Indonesian STEM textbook data, containing 5,169 books and 208.30 million words, designed to support the development and training of advanced NLP systems and AI models for scientific understanding, problem-solving, and concept learning in Bahasa. Full Dataset Overview This dataset is part of a large-scale multilingual educational corpus containing over 3+ billion words across 5,000+ subjects, supported by interwoven images for… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Indonesian-STEM-Textbook-Dataset.tabular10K<n<100K0 likes35 downloads2d agoHugging Face22InfoBayAI /English-STEM-Textbook-DatasetgatedDataset Description: This dataset is a large-scale collection of English STEM textbook data, containing 8,939 books and 855.01 million words, designed to support the development and training of advanced NLP systems and AI models for scientific understanding, problem-solving, and concept learning in English. Full Dataset Overview This dataset is part of a large-scale multilingual educational corpus containing over 3+ billion words across 5,000+ subjects, supported by interwoven images for… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/English-STEM-Textbook-Dataset.tabular100K<n<1M0 likes34 downloads2d agoHugging Face23fidaakh /STEM_datatabular10K<n<100K1 likes32 downloads2y agoHugging Face24electricsheepafrica /africa-ilo-emp-stem-sex-eco-nb-employment-in-stem-occupations-by-sex-and-economic Employment in STEM occupations by sex and economic activity (thousands) | Africa (ILOSTAT) | Africa (Electric Sheep Africa metadata inventory) Size category: 1K<n<10K - Formats: parquet - Sector: economics_finance - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-ilo-emp-stem-sex-eco-nb-employment-in-stem-occupations-by-sex-and-economic.tabulartabular-classification1K<n<10K0 likes30 downloads2mo agoHugging Face25InfoBayAI /Javanese-Non-STEM-Textbook-DatasetgatedDataset Description: This dataset is a large-scale collection of Javanese Non-STEM textbook data, containing 597 books and 67.05 million words, designed to support the development and training of advanced NLP systems and AI models for language understanding, reasoning, and general knowledge learning in Javanese. Full Dataset Overview This dataset is part of a large-scale multilingual educational corpus containing over 3+ billion words across 5,000+ subjects, supported by interwoven images for… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Javanese-Non-STEM-Textbook-Dataset.tabular10K<n<100K0 likes26 downloads2d agoHugging Face26InfoBayAI /Indonesian-Non-STEM-Textbook-DatasetgatedDataset Description: This dataset is a large-scale collection of Indonesian Non-STEM textbook data, containing 4,098 books and 182.10 million words, designed to support the development and training of advanced NLP systems and AI models for language understanding, reasoning, and general knowledge learning in Bahasa. Full Dataset Overview This dataset is part of a large-scale multilingual educational corpus containing over 3+ billion words across 5,000+ subjects, supported by interwoven images… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Indonesian-Non-STEM-Textbook-Dataset.tabular10K<n<100K0 likes25 downloads2d agoHugging Face27electricsheepafrica /africa-ilo-emp-stem-sex-ste-nb-employment-in-stem-occupations-by-sex-and-status-i Employment in STEM occupations by sex and status in employment (thousands) | Africa (ILOSTAT) | Africa (Electric Sheep Africa metadata inventory) Size category: 1K<n<10K - Formats: parquet - Sector: economics_finance - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-ilo-emp-stem-sex-ste-nb-employment-in-stem-occupations-by-sex-and-status-i.tabulartabular-classification1K<n<10K0 likes25 downloads2mo agoHugging Face28InfoBayAI /Marathi-STEM-Textbook-DatasetgatedDataset Description: This dataset is a large-scale collection of Marathi STEM textbook data, containing 173 books and 7.81 million words, designed to support the development and training of advanced NLP systems and AI models for scientific understanding, problem-solving, and concept learning in Marathi. Full Dataset Overview This dataset is part of a large-scale multilingual educational corpus containing over 3+ billion words across 5,000+ subjects, supported by interwoven images for deeper… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Marathi-STEM-Textbook-Dataset.tabular10K<n<100K0 likes25 downloads2d agoHugging Face29InfoBayAI /Arabic-STEM-Textbook-DatasetgatedDataset Description: This dataset is a large-scale collection of Arabic STEM textbook data, containing 1,364 books and 63.51 million words, designed to support the development and training of advanced NLP systems and AI models for scientific understanding, problem-solving, and concept learning in Arabic. Full Dataset Overview This dataset is part of a large-scale multilingual educational corpus containing over 3+ billion words across 5,000+ subjects, supported by interwoven images for deeper… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Arabic-STEM-Textbook-Dataset.tabulartext-classification10K<n<100K0 likes24 downloads2d agoHugging Face30InfoBayAI /Bengali-STEM-Textbook-DatasetgatedDataset Description: This dataset is a large-scale collection of Bengali STEM textbook data, containing 308 books and 12.88 million words, designed to support the development and training of advanced NLP systems and AI models for scientific understanding, problem-solving, and concept learning in Bengali. Full Dataset Overview This dataset is part of a large-scale multilingual educational corpus containing over 3+ billion words across 5,000+ subjects, supported by interwoven images for deeper… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Bengali-STEM-Textbook-Dataset.tabular10K<n<100K0 likes24 downloads2d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.