datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
United_States_State_Legislation_with_SummariesTest Push
arxiv-software-repo-links-datacite-enrichment-format
arXiv Software Repository Links - DataCite Enrichment Format
A collection of metadata enrichments, formatted for DataCite's enrichment API, that add links between arXiv papers (via DOI) and the software repositories they reference or are supplemented by.
Quick Start
from datasets import load_dataset
ds = load_dataset("cometadata/arxiv-software-repo-links-datacite-enrichment-format")
Dataset Description
Each record is a DataCite-style enrichment instruction… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/arxiv-software-repo-links-datacite-enrichment-format.arxiv-software-repo-links
arXiv Software Repository Links
A dataset mapping arXiv papers (via DOI) to software repositories they reference or which are implementations of the work. Additionaly includes co-citation analysis and community clustering.
Implementation detection is conducted using the evamxb/dev-author-em-clf model from sci-soft-models
Quick Start
from datasets import load_dataset
# Load DOI-to-repo links
links = load_dataset("cometadata/arxiv-software-repo-links", "links")
#… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/arxiv-software-repo-links.k12-standards-instruction-tasks
K-12 Curriculum Tasks (generated)
2,489 generated instruction/input/output records covering five curriculum tasks:
assessment creation, learning objective generation, misconception detection, standard
explanation, and standards Q&A. Content is predominantly mathematics.
Important: the name is misleading
Despite the name, this dataset contains no school directory data. There are four
columns - task, input, output, metadata - and no staff, principal, or school… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/k12-standards-instruction-tasks.software-strategist-v1
Software Fundamentals — Strategy Knowledge Base
A language-agnostic knowledge base of software engineering fundamentals, paired with a synthetic instruction-tuning dataset (~13,500 examples) for training small language models (SLMs) as software engineering strategists.
The trained model takes a description of a coding situation and routes it to relevant concepts, outputting synthesized strategic guidance as structured JSON.
Dataset Summary
This dataset provides ~13… See the full description on the dataset page: https://huggingface.co/datasets/jtregunna/software-strategist-v1.texas-k12-curriculum-standards-teks
Texas K-12 Curriculum Standards (TEKS-derived)
15,040 generated learning-objective records organized around the Texas Essential
Knowledge and Skills (TEKS) taxonomy, spanning core academic subjects, Career & Technical
Education clusters, and specialized program areas.
How this was built (read this first)
These records are programmatically generated, not transcribed from official standards
documents. A generator took a standards taxonomy - codes, grade levels… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/texas-k12-curriculum-standards-teks.file-format-software-version-compatibility
GIS vector format capabilities and limitations in GDAL
Canonical, always-current version: https://referencesource.org/file-format-software-version-compatibility/
Machine-readable: https://referencesource.org/file-format-software-version-compatibility/data.json — this mirror is a point-in-time copy.
Last verified: 2026-10-06
Stale after: 2027-01-31 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 3
Key capabilities… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/file-format-software-version-compatibility.screenplay-format-edge-cases
Screenplay Format Edge Cases
48 original Fountain specimens in 24 contrast pairs — two near-identical inputs
per pair, at the points where the Fountain syntax leaves a choice. In 17 pairs
the one difference changes how the lines are classified. In the other 7 it
changes the surface and the labels hold: a lowercase scene prefix, a cue
extension, a non-Latin cue, escaped characters, a dual-dialogue caret, an inline
note and centered-text markers.
Version: 1.0.0 · Maintainer:… See the full description on the dataset page: https://huggingface.co/datasets/Inkwell-Software/screenplay-format-edge-cases.software-eol-change-feed
Software end-of-life date changes — which dates moved and what they were before
Canonical, always-current version: https://referencesource.org/software-eol-change-feed/
Machine-readable: https://referencesource.org/software-eol-change-feed/data.json — this mirror is a point-in-time copy.
Last verified: 2026-10-05
Stale after: 2026-10-19 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 1999
A change log of software… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/software-eol-change-feed.nanoset
Sourceworks NanoSet
NanoSet is an experimental dataset where the main goal is to create a usable chatbot through less training data.
What is in NanoSet?
NanoSet is divded into 3 major sections, containg 36 entries divided into 6 sub-topics. The structure creates 108 total lines of training data, which may be subject to change in the future. The following is a visual on the structure:
108 entries total
3 Sections, each with 36 entries:
Chat Basics (Greetings, Jokes, etc.)… See the full description on the dataset page: https://huggingface.co/datasets/srcworks-software/nanoset.screenplay-revision-evaluation
Screenplay Revision Evaluation Cases
24 original screenwriting revision tasks. Each gives a short scene and a
constraint — cut a page to its beat, plant a prop, hold an answer back, fix a
continuity slip — then pairs it with mechanical checks (a word ceiling, a line
that must survive) and separate human-review questions. It tests whether a tool,
or a person, can make a tightly-constrained edit while keeping the scene intact.
Each task's reference_output is null, because a… See the full description on the dataset page: https://huggingface.co/datasets/Inkwell-Software/screenplay-revision-evaluation.software-end-of-support
Software end-of-support dates
Canonical, always-current version: https://referencesource.org/software-end-of-support/
Machine-readable: https://referencesource.org/software-end-of-support/data.json — this mirror is a point-in-time copy.
Last verified: 2026-09-30
Stale after: 2026-12-29 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 61
Release, end-of-active-support, end-of-life (EOL) and end-of-security-support dates… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/software-end-of-support.software-interface-version-compatibility
Software interface version compatibility
Canonical, always-current version: https://referencesource.org/software-interface-version-compatibility/
Machine-readable: https://referencesource.org/software-interface-version-compatibility/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-04
Stale after: 2026-10-03 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 417
Which versions of common software… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/software-interface-version-compatibility.mimo-openenv-software
MiMo curriculum for OpenEnv
Twelve original tasks from XiaomiMiMo/MiMo-V2.6-RL-oss, served through one OpenEnv shell/finish interface. Nine are code tasks and three are terminal tasks. Problem statements, test patches, embedded test files and scoring rules are preserved. No LLM judge or external service credentials are required.
Source repository · Public dataset
Curriculum
tasks.jsonl records task identities, immutable upstream images, source positions and… See the full description on the dataset page: https://huggingface.co/datasets/akseljoonas/mimo-openenv-software.Software-Architectural-FrameworksSoftware-Architectural-Frameworks
I am releasing a small dataset covering topics related to Frameworks under Software-Architecture.
I have included following topics:
TOGAF
Zachman Framework
IEEE 1471
Matrix-based approach to architecture development
Significance of IEEE 1471 (ISO/IEC 42010)
Benefits of employing architectural frameworks
and Many More!
This dataset can be useful in LLM development. Also those who are working on developing Software development related LLMs then this dataset can… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/Software-Architectural-Frameworks.agentic-software-conformance
TeaQL Agentic Software Conformance
Machine-readable evidence for the TeaQL Harness: semantic-model evaluation,
generated artifacts, seven language-native runtimes, executable examples, and
cross-language conformance checks.
This is an evidence dataset, not a leaderboard and not a collection of
unverified model claims. Each row identifies its evidence level, exact source,
verification date, revisions where available, command or gate, result, and
important qualifications. The… See the full description on the dataset page: https://huggingface.co/datasets/teaql/agentic-software-conformance.ontocord__wide_3b_sft_stag1.2-lyrical_news_software_howto_formattedtext-merge-details
Dataset Card for Evaluation run of ontocord/wide_3b_sft_stag1.2-lyrical_news_software_howto_formattedtext-merge
Dataset automatically created during the evaluation run of model ontocord/wide_3b_sft_stag1.2-lyrical_news_software_howto_formattedtext-merge
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ontocord__wide_3b_sft_stag1.2-lyrical_news_software_howto_formattedtext-merge-details.arxiv-software-repo-links
arXiv Software Repository Links
A dataset mapping arXiv papers (via DOI) to software repositories they reference or which are implementations of the work. Additionaly includes co-citation analysis and community clustering.
Implementation detection is conducted using the evamxb/dev-author-em-clf model from sci-soft-models
Quick Start
from datasets import load_dataset
# Load DOI-to-repo links
links = load_dataset("cometadata/arxiv-software-repo-links", "links")
#… See the full description on the dataset page: https://huggingface.co/datasets/rafidirtiza/arxiv-software-repo-links.real-toxicity-prompts
Dataset Card for Real Toxicity Prompts
Dataset Summary
RealToxicityPrompts is a dataset of 100k sentence snippets from the web for researchers to further address the risk of neural toxic degeneration in models.
Languages
English
Dataset Structure
Data Instances
Each instance represents a prompt and its metadata:
{
"filename":"0766186-bc7f2a64cb271f5f56cf6f25570cd9ed.txt",
"begin":340,
"end":564,
"challenging":false… See the full description on the dataset page: https://huggingface.co/datasets/ravi-softwarethreads/real-toxicity-prompts.software-testing-ai-agent
Software Testing Agent Meta and Traffic Dataset in AI Agent Marketplace | AI Agent Directory | AI Agent Index from DeepNLP
This dataset is collected from AI Agent Marketplace Index and Directory at http://www.deepnlp.org, which contains AI Agents's meta information such as agent's name, website, description, as well as the monthly updated Web performance metrics, including Google,Bing average search ranking positions, Github Stars, Arxiv References, etc.
The dataset is helpful for… See the full description on the dataset page: https://huggingface.co/datasets/DeepNLP/software-testing-ai-agent.ontocord__wide_3b_sft_stage1.2-ss1-expert_software-details
Dataset Card for Evaluation run of ontocord/wide_3b_sft_stage1.2-ss1-expert_software
Dataset automatically created during the evaluation run of model ontocord/wide_3b_sft_stage1.2-ss1-expert_software
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ontocord__wide_3b_sft_stage1.2-ss1-expert_software-details.software-engineering-and-devopsecho-software-loraAI_software_datasetenglish-software-engineering-basics-30software-testing-datasettranslated_dataset_software_Questions
