datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Python-Code-Solutions
Python Code Solutions
Features
1000k of Python Code Solutions for Text Generation and Question Answering
Python Coding Problems labelled by topic and difficulty
Recommendations
Train your Model on Logical Operations and Mathematical Problems Before Training it on this. This is optional for Fine Tuning 2B parameter + models.
Format the prompts in a orderly way when formatting data eg. {question} Solution: {solution} Topic: {topic}
software_requirementsk12-standards-instruction-tasks
K-12 Curriculum Tasks (generated)
2,489 generated instruction/input/output records covering five curriculum tasks:
assessment creation, learning objective generation, misconception detection, standard
explanation, and standards Q&A. Content is predominantly mathematics.
Important: the name is misleading
Despite the name, this dataset contains no school directory data. There are four
columns - task, input, output, metadata - and no staff, principal, or school… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/k12-standards-instruction-tasks.software-strategist-v1
Software Fundamentals — Strategy Knowledge Base
A language-agnostic knowledge base of software engineering fundamentals, paired with a synthetic instruction-tuning dataset (~13,500 examples) for training small language models (SLMs) as software engineering strategists.
The trained model takes a description of a coding situation and routes it to relevant concepts, outputting synthesized strategic guidance as structured JSON.
Dataset Summary
This dataset provides ~13… See the full description on the dataset page: https://huggingface.co/datasets/jtregunna/software-strategist-v1.texas-k12-curriculum-standards-teks
Texas K-12 Curriculum Standards (TEKS-derived)
15,040 generated learning-objective records organized around the Texas Essential
Knowledge and Skills (TEKS) taxonomy, spanning core academic subjects, Career & Technical
Education clusters, and specialized program areas.
How this was built (read this first)
These records are programmatically generated, not transcribed from official standards
documents. A generator took a standards taxonomy - codes, grade levels… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/texas-k12-curriculum-standards-teks.jeopardy-clues
Jeopardy! Clues
568,068 Jeopardy! clues with their answers, categories, dollar values, air dates, and
round information, compiled from publicly archived, community-maintained transcriptions
of aired episodes.
Loading
from datasets import load_dataset
ds = load_dataset("robworks-software/jeopardy-clues")
science = ds["train"].filter(lambda x: x["category"] == "SCIENCE")
Splits
Split
Rows
train
482,857
validation
42,605
test
42,606… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/jeopardy-clues.cosimo-cfa-frm-71k
Cosimo: Synthetic CFA/FRM Financial Reasoning Dataset
Cosimo is a synthetic, code-verified financial-exam question dataset for
training reasoning models and preference-tuned (DPO/ORPO) models. It contains
71,000 original, numerically-grounded questions spanning the CFA Level I–III
and FRM Part 1/2 curricula, each with a step-by-step chain-of-thought reasoning
trace.
Every numerical answer is computed by reference code, never sampled from a
language model. Reasoning traces are… See the full description on the dataset page: https://huggingface.co/datasets/btech-software/cosimo-cfa-frm-71k.SwiftUI-Code-Examples
SwiftUI Code Solutions
Dataset Created by MCES10 Software has SwiftUI Code Problems and can be used for AI training for Code Generation
Recommendations
Train your LLM on the Swift and SwiftUI Framework Syntax before training it this
Fine Tune or Train Effectively at optimal Epochs and Learning Rates
Use the whole dataset for training
Your Model may need to be Prompt Tuned for the best performance but it isn't required.
Use test when testing or trialing the dataset
Use… See the full description on the dataset page: https://huggingface.co/datasets/MCES10-Software/SwiftUI-Code-Examples.nanoset
Sourceworks NanoSet
NanoSet is an experimental dataset where the main goal is to create a usable chatbot through less training data.
What is in NanoSet?
NanoSet is divded into 3 major sections, containg 36 entries divided into 6 sub-topics. The structure creates 108 total lines of training data, which may be subject to change in the future. The following is a visual on the structure:
108 entries total
3 Sections, each with 36 entries:
Chat Basics (Greetings, Jokes, etc.)… See the full description on the dataset page: https://huggingface.co/datasets/srcworks-software/nanoset.screenplay-revision-evaluation
Screenplay Revision Evaluation Cases
24 original screenwriting revision tasks. Each gives a short scene and a
constraint — cut a page to its beat, plant a prop, hold an answer back, fix a
continuity slip — then pairs it with mechanical checks (a word ceiling, a line
that must survive) and separate human-review questions. It tests whether a tool,
or a person, can make a tightly-constrained edit while keeping the scene intact.
Each task's reference_output is null, because a… See the full description on the dataset page: https://huggingface.co/datasets/Inkwell-Software/screenplay-revision-evaluation.k12-ela-standards-expanded
K-12 ELA Standards, expanded (generated instruction data)
12,282 instruction/input/output records for English Language Arts, generated around a K-12
standards taxonomy for instruction-tuning and educational-content experiments.
How this was built (read this first)
These are programmatically generated training examples, not curriculum written by
educators and not the text of any official standard. A generator combined standards
metadata - codes, grade levels… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/k12-ela-standards-expanded.database-query-logs-synthetic
Database Query Logs (synthetic)
3,995 database query-log entries spanning 10 engines - MySQL, PostgreSQL, MongoDB, SQL
Server, Oracle, MariaDB, SQLite, Cassandra, Redis, and Elasticsearch - with query text,
type, complexity, execution timing, and row-count metadata.
These queries are synthetic
The queries were programmatically generated, not captured from production systems.
They were produced by templating a set of query shapes across industry-flavored schema… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/database-query-logs-synthetic.k12-science-standards
[!WARNING]
Deprecated - use k12-science-standards-expanded instead.
This dataset is superseded: every instruction in this set also appears there, plus 1,123 more and nine additional metadata columns. Nothing here is unique to it.
It stays online so existing references keep resolving, but it will not be updated.
New work should point at robworks-software/k12-science-standards-expanded.
K-12 Science Standards (generated instruction data)
6,787 instruction/input/output records… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/k12-science-standards.k12-science-standards-expanded
K-12 Science Standards, expanded (generated instruction data)
15,354 instruction/input/output records for science, generated around a K-12
standards taxonomy for instruction-tuning and educational-content experiments.
How this was built (read this first)
These are programmatically generated training examples, not curriculum written by
educators and not the text of any official standard. A generator combined standards
metadata - codes, grade levels, domains… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/k12-science-standards-expanded.ccisd-teks-training
[!WARNING]
Deprecated - use ccisd-teks-enhanced instead.
This dataset is superseded: both cover the same 3,628 inputs, but that one carries eight further columns (teaching strategies, misconceptions, assessment examples and more). Nothing here is unique to it.
It stays online so existing references keep resolving, but it will not be updated.
New work should point at robworks-software/ccisd-teks-enhanced.
CCISD TEKS Training Set (generated)
4,224 instruction-tuning examples… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/ccisd-teks-training.CPP-Code-Solutions
C++ Code Solutions
Features
1000k of Python Code Solutions for Text Generation and Question Answering
C++ Coding Problems labelled by topic and difficulty
Recommendations
Train your Model on Logical Operations and Mathematical Problems Before Training it on this. This is optional for Fine Tuning 2B parameter + models.
Format the prompts in a orderly way when formatting data eg. {question} Solution: {solution} Topic: {topic}
k12-mathematics-standards-aligned
[!WARNING]
Deprecated - use k12-mathematics-standards-expanded instead.
This dataset is superseded: every input in this set also appears there, plus 366 more and two additional metadata columns. Nothing here is unique to it.
It stays online so existing references keep resolving, but it will not be updated.
New work should point at robworks-software/k12-mathematics-standards-expanded.
K-12 Mathematics Standards (generated instruction data)
4,397 instruction/input/output records… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/k12-mathematics-standards-aligned.historical-training-manuals
Historical Training Manuals
1,597 US government and government-adjacent training manuals and technical publications
sourced from the Internet Archive, spanning roughly 1800-2021. Records carry
bibliographic metadata; a subset also carries extracted full text and a machine-generated
summary.
Loading
from datasets import load_dataset
ds = load_dataset("robworks-software/historical-training-manuals")
Splits
Split
Rows
train
1,277… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/historical-training-manuals.IFEval-Multi-IF-hy
IFEval-Multi-IF-hy — Armenian IFEval & Multi-IF
Dataset Summary
We introduce IFEval & Multi-IF hy, an Armenian extension of Multi-IF, the benchmark for assessing LLMs' proficiency in following multi-turn and multilingual instructions. Multi-IF itself extends IFEval to multi-turn, multilingual conversations. Multi-IF covers eight languages and does not include Armenian. To build IFEval & Multi-IF hy, the English conversations were first split into two groups:… See the full description on the dataset page: https://huggingface.co/datasets/Center-Of-Advanced-Software-Technologies/IFEval-Multi-IF-hy.k12-business-economics-standards
K-12 Business and Economics Standards
1,236 generated learning-objective records covering financial literacy, personal finance,
entrepreneurship, business management, and career development, organized around the
Jump$tart Personal Financial Education and NBEA Business Education standard structures.
How this was built (read this first)
These records are programmatically generated, not transcribed from official standards
documents. A generator took a standards… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/k12-business-economics-standards.k12-ela-standards
[!WARNING]
Deprecated - use k12-ela-standards-expanded instead.
This dataset is superseded: every input in this set also appears there, plus 1,433 more and five additional metadata columns. Nothing here is unique to it.
It stays online so existing references keep resolving, but it will not be updated.
New work should point at robworks-software/k12-ela-standards-expanded.
K-12 ELA Standards (generated instruction data)
6,487 instruction/input/output records for English Language… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/k12-ela-standards.k12-mathematics-standards-expanded
K-12 Mathematics Standards, expanded (generated instruction data)
4,965 instruction/input/output records for mathematics, generated around a K-12
standards taxonomy for instruction-tuning and educational-content experiments.
How this was built (read this first)
These are programmatically generated training examples, not curriculum written by
educators and not the text of any official standard. A generator combined standards
metadata - codes, grade levels, domains… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/k12-mathematics-standards-expanded.k12-social-studies-standards
K-12 Social Studies Standards (generated instruction data)
15,982 instruction/input/output records for social studies (civics, history, geography, economics), generated around a K-12
standards taxonomy for instruction-tuning and educational-content experiments.
How this was built (read this first)
These are programmatically generated training examples, not curriculum written by
educators and not the text of any official standard. A generator combined standards… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/k12-social-studies-standards.california-k12-standards
California K-12 Educational Standards
3,410 records organized around California K-12 standards frameworks, including Common
Core, NGSS, ELD, CTE, and Ethnic Studies. Records carry a standard identifier, grade
level, subject area, domain, and generated learning-objective and application text.
How this was built (read this first)
These records are programmatically generated, not transcribed from official standards
documents. A generator took a standards taxonomy -… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/california-k12-standards.ccisd-teks-enhanced
CCISD TEKS Enhanced (LLM-generated)
4,224 records built from the same 428 TEKS expectations as
ccisd-teks-training,
with additional LLM-written fields: detailed explanations, real-world applications,
prerequisite knowledge, common misconceptions, teaching strategies, assessment examples,
cross-curricular connections, and learning progressions.
The added content is LLM output and was not reviewed
The enrichment fields were generated by a language model. No educator… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/ccisd-teks-enhanced.JS-Code-Solutions
Python Code Solutions
Features
1000k of JS Code Solutions for Text Generation and Question Answering
JS Coding Problems labelled by topic and difficulty
Recommendations
Train your Model on Logical Operations and Mathematical Problems Before Training it on this. This is optional for Fine Tuning 2B parameter + models.
Format the prompts in a orderly way when formatting data eg. {question} Solution: {solution} Topic: {topic}
agentic-software-conformance
TeaQL Agentic Software Conformance
Machine-readable evidence for the TeaQL Harness: semantic-model evaluation,
generated artifacts, seven language-native runtimes, executable examples, and
cross-language conformance checks.
This is an evidence dataset, not a leaderboard and not a collection of
unverified model claims. Each row identifies its evidence level, exact source,
verification date, revisions where available, command or gate, result, and
important qualifications. The… See the full description on the dataset page: https://huggingface.co/datasets/teaql/agentic-software-conformance.ccisd-unified-master-2024
CCISD Unified School Master (2024)
School-level records for Clear Creek Independent School District (Texas), compiled from
the district's public school pages and Texas Education Agency accountability reports.
Covers 39 schools with principal names, contact details, enrollment, and accountability
ratings.
Loading
from datasets import load_dataset
ds = load_dataset("robworks-software/ccisd-unified-master-2024")
all_schools = ds["full"] # all 39 schools… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/ccisd-unified-master-2024.ccisd-teks-alignment-split
[!WARNING]
Deprecated - use ccisd-teks-alignment instead.
This dataset is superseded: the two contain the same 428 rows with the same 12 columns; this copy only adds a train/validation/test partition, which you can reproduce in one line. Nothing here is unique to it.
It stays online so existing references keep resolving, but it will not be updated.
New work should point at robworks-software/ccisd-teks-alignment.
CCISD TEKS Alignment (pre-split)
The same 428 TEKS-to-course… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/ccisd-teks-alignment-split.ccisd-teks-alignment
CCISD TEKS Alignment
428 Texas Essential Knowledge and Skills (TEKS) student expectations mapped to 25 Clear
Creek ISD high school courses, with STAAR-tested status flagged.
Loading
from datasets import load_dataset
ds = load_dataset("robworks-software/ccisd-teks-alignment")
428 rows, single train split. A pre-split version of the same 428 rows is published as
ccisd-teks-alignment-split.
Contents
428 distinct TEKS codes (e.g. ELAR.9.1.A), each… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/ccisd-teks-alignment.
