datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arabic-stem-lexicon
Arabic Diacritized-Stem Lexicon
An undiacritized Arabic surface form → its most frequent diacritized stem.
Standard Arabic writes no short vowels, so anything that has to pronounce Arabic
must first put them back. A neural diacritizer does that well on rare words, where
inference is the only thing there is. On common words it is the wrong tool:
which vowels كتاب carries is not a thing to be inferred, it is a thing to be looked
up — and models get exactly these wrong, reading… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/arabic-stem-lexicon.stem-diagrams
STEM Diagrams
30,325 technical diagrams (block diagrams, schematics, flowcharts, architectures)
extracted from arXiv papers across six engineering fields, each with a source
attribution and a quality score. Built by an LLM-curated pipeline and used to show
that a small frozen-feature classifier can replace the paid LLM labeling gate.
Paper: Distilling an LLM Diagram-Curation Pipeline into Local Classifiers (Adnan Abbasi, Thothica, 2026)
Code:… See the full description on the dataset page: https://huggingface.co/datasets/aeyxen/stem-diagrams.Biomimicry-Nectar-BioDesign-STEMCreated by combining the strongest parts of the ANIMA training datasets. RAG was used to further enhance and factuality check as much as possible.
bazi-hidden-stemsfrom datasets import load_dataset
rows = load_dataset("Shann5/bazi-hidden-stems") # one row per (source, branch, stem)
compare = load_dataset("Shann5/bazi-hidden-stems", "compare") # one row per branch, one column per source
Mirror of hidden-stems/ in
Shann5/bazi-open-data — the GitHub copy is canonical
(JSON version, source texts, build). Corrections belong there as issues.
Hidden stems (藏干) across sources
Every earthly branch stores one to three heavenly… See the full description on the dataset page: https://huggingface.co/datasets/Shann5/bazi-hidden-stems.Pop-Rock-Hybrid-Stem-Dataset-cat001
Dataset Overview: Pop Rock Hybrid Stem Dataset (cat001)
This dataset contains a curated collection of original instrumental music designed for commercial and research applications in music analysis, audio modeling, and production workflows.
Every composition, arrangement, performance, sound design element, and production decision was created entirely through human musical and technical processes. All music contained in this dataset is 100% human-made (is_human_created: TRUE).… See the full description on the dataset page: https://huggingface.co/datasets/ToneCubeMedia/Pop-Rock-Hybrid-Stem-Dataset-cat001.Hindi-Non-STEM-QA-MCQ-DatasetDataset Description:
This dataset is a large-scale collection of Hindi Non-STEM Question Answering (QA) data, containing over 1.4 million question-answer pairs, designed to support the development and training of advanced NLP systems and AI models for language understanding, reasoning, knowledge retrieval, and educational learning in Hindi. It is part of a broader collection of 6.5+ million question-answer pairs spanning multiple languages and domains.
The dataset consists of multiple-choice… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Hindi-Non-STEM-QA-MCQ-Dataset.STEM_datastem-programmes-top-100-universities-2025
STEM Programmes at the Top 100 Universities, 2025
Which STEM degrees do the world's 100 highest-ranked universities actually teach? 1,258 university × programme pairs, with the synonym table used to make programme names comparable.
Method, tables and the essays that use this data: https://mariascales.com/research/long-bet/data/programmes/
What is in here
stem_programmes_arwu100_2025.csv – 1,258 rows. Columns: rank (ARWU 2025, 1–100, ties kept), university… See the full description on the dataset page: https://huggingface.co/datasets/MariaScales/stem-programmes-top-100-universities-2025.Wiki_STEM_Corpusafrica-synth-education-stem-education-access-africa-all
Africa Synth Education Stem Education Access Africa All | Africa (World Bank)
Size category: 10K<n<100K - Formats: csv - Sector: education - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Education datasets help… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-education-stem-education-access-africa-all.Turkish-STEM-DPO-Dataset
Not: Bu veri setinin dokümantasyonu Türk yapay zeka topluluğuna katkı sağlamak amacıyla VeriPazarı tarafından Türkçeye çevrilmiştir. Orijinal veri seti yusufbaykaloglu tarafından geliştirilmiş olup, VeriPazarı tarafından Türk AI ekosistemi için arşivlenmiştir.
🔗 Orijinal Kaynak: yusufbaykaloglu/Turkish-STEM-DPO-Dataset
🔗 Derleyen Platform: VeriPazarı
Türkçe STEM DPO Veri Seti (Turkish STEM DPO Dataset)
Veri Seti Özeti
Turkish STEM DPO (Doğrudan Tercih… See the full description on the dataset page: https://huggingface.co/datasets/Taklaxbr/Turkish-STEM-DPO-Dataset.somali-stem-dataset
Somali STEM Terminology Dataset (Sample)
Overview
A bilingual English–Somali terminology dataset covering 5 STEM domains.
This is a 100-row public sample. The full dataset contains 3,000+ terms.
Somali is spoken by 20+ million people but is critically under-represented in AI training data. This is the only known structured Somali STEM lexicon in machine-readable format.
Fields
Field
Description
English
Scientific term in English
Somali
Somali… See the full description on the dataset page: https://huggingface.co/datasets/planwise-data/somali-stem-dataset.Turkish-STEM-DPO-Dataset
Turkish STEM DPO Dataset
Dataset Summary
The Turkish STEM DPO (Direct Preference Optimization) dataset is a comprehensive synthetic resource containing 16,177 high-quality preference pairs designed to enhance the reasoning capabilities of Turkish language models in mathematics, physics, and programming.
The dataset leverages a preference-based learning approach: each instance pairs a carefully crafted, expert-level solution with a deliberately flawed or incomplete… See the full description on the dataset page: https://huggingface.co/datasets/yusufbaykaloglu/Turkish-STEM-DPO-Dataset.STEM
Standards-Targeted Educational Math (STEM) Dataset
The STEM dataset contains 2,577 high-quality math word problems for 3rd-5th grade students either annotated by teachers and Gemma 3 27B IT or from the 3rd-5th grade subset of ASDIV, as reported in EDUMATH: Generating Standards-aligned Educational Math Word Problems. Each row contains a question and solution along with its associated grade level and math standard from the Virginia Standards of Learning. For further information about… See the full description on the dataset page: https://huggingface.co/datasets/bryanchrist/STEM.WIKI-STEM-CORPUSnews_stemm_esNews spanish media outlets
H4StackExchange_STEM_small_answerssmall_wiki_stem_chunked_mcqastem_csvRAG-STEM-Wikilinkbricks_ko_dataset_stem_2original-stem-practice-bank
🎓 Original STEM Practice Bank
A collection of original, self-authored questions for Computer Science, Physics, and Biology.
linkbricks_ko_dataset_stem_1Wiki_STEM_Corpusstem_csv_1stem-rag-corpussmall_wiki_stem_chunkedwiki-stem-csv-1sigma-gennaiosmall_wiki_stem
