datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VOE-Bench
VOE-Bench 2.2 Core
VOE-Bench asks whether an agent can tell when the evidence in an archived scientific workflow is enough to act—and when it should keep reading, escalate, refuse, or identify a broken source record.
The benchmark asks a narrow question: given a frozen workflow archive and an explicit evidence budget, can an agent acquire the right records, track their provenance, and stop for the right reason?
Version 2.2 Core is a 91-task public development and… See the full description on the dataset page: https://huggingface.co/datasets/Dynamical-Systems/VOE-Bench.systems_programming_and_administrationclassical-cipher-corpus
Classical Cipher Corpus
A labeled educational dataset of classical cipher examples for teaching cryptanalysis and training small cipher-family classifiers.
Part of the Cipher Detective AI project:
🕵️ Space: systemslibrarian/cipher-detective-ai
📦 Dataset: systemslibrarian/classical-cipher-corpus (this repo)
🤖 Model: systemslibrarian/cipher-detective-classifier
Intended use
Teach classical cryptanalysis.
Benchmark educational cipher-family detectors.
Train small… See the full description on the dataset page: https://huggingface.co/datasets/systemslibrarian/classical-cipher-corpus.ai-system-patterns
AI System Patterns
A compact reference dataset of reusable architectural patterns for modern AI systems.
The dataset focuses on practical system-design concepts across AI agents, orchestration, memory, validation, observability, interoperability, world models, Physical AI, data pipelines, and production operations.
Each row contains:
pattern
category
description
components
use_case
complexity
Example
{
"pattern": "Model Routing",
"category": "Orchestration"… See the full description on the dataset page: https://huggingface.co/datasets/ai-systems/ai-system-patterns.repro-learning-to-share-selective-memory-for-efficient-parallel-agentic-systems-traces
Agent traces
Agent sessions published from a Trackio Logbook.
ml-systems-interview-bench
ML Systems Interview Bench
ML Systems Interview Bench is a structured, benchmark-style dataset for evaluating technical interview answers across practical ML engineering, MLOps, model serving, ML system design, data pipelines, LLM/RAG, observability, Python engineering, and production debugging.
Each record combines an interview-style question with expected concepts, a concise reference answer, qualitative evaluation anchors, question-specific skills, and follow-up questions.… See the full description on the dataset page: https://huggingface.co/datasets/Max00035/ml-systems-interview-bench.library-classification-systems
Library Classification Systems
This comprehensive dataset contains hierarchical outlines of major library classification systems, offering a valuable resource for researchers, librarians, and information scientists.
Classification System
Abbreviation
Primary Usage
Language
Entries
Dewey Decimal Classification
DDC
International
English
1110
Library of Congress Classification
LCC
International
English
6517
Universal Decimal Classification
UDC
International
English
2431… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/library-classification-systems.agentic-systems-showcase
Agentic Systems Showcase
Catalog of public projects by Hanumanthu Harsha Vardhan
(GitHub hharsha98, Hub hharsha).
Each row is an honest summary taken from the project's own README, Space card, or studio site.
No invented papers or benchmarks.
Field
Meaning
id
Stable slug
name
Display name
kind
github, space, or website
url
Canonical link
summary
One-paragraph description
stack
Coarse tech/topic tags
owner
GitHub or Hub handle
Rows: 16 (expanded from… See the full description on the dataset page: https://huggingface.co/datasets/hharsha/agentic-systems-showcase.repro-latent-collaboration-in-multi-agent-systems-traces
Agent traces
Agent sessions published from a Trackio Logbook.
embedded-systems-qa
Embedded Systems Engineering Q&A — Instruction Dataset
A hand-authored instruction-tuning dataset of technical question/answer pairs for
embedded systems engineering, formatted for supervised fine-tuning of Mistral 7B
(Alpaca-style instruction / input / output schema).
At a glance
Entries
302
Format
JSONL, one JSON object per line
Schema
{"instruction": <question>, "input": "", "output": <answer>}
Language
English
Avg. answer length
~590… See the full description on the dataset page: https://huggingface.co/datasets/eniomecaj/embedded-systems-qa.systems_programming_code_conversationsrag-systems-sft-100k
RAG Systems SFT 100K
A synthetic supervised fine-tuning dataset of 100,000 high-quality conversations covering Retrieval-Augmented Generation (RAG) systems — from basic pipelines to advanced multi-hop retrieval, evaluation, and production optimization. Designed to train AI assistants that can help engineers build, debug, and scale RAG applications.
Dataset Description
This dataset covers the full spectrum of RAG system development across 12 specialized categories.… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/rag-systems-sft-100k.principal-systems-architect-dataset
Principal Systems Architect & Kernel Engineering Dataset
A high-quality, expert-level synthetic dataset of 200 comprehensive scenarios, questions, rationales, and detailed solutions focused on High-Performance Distributed Systems, Kernel Architecture, and Systems Programming.
Each entry has been carefully structured, validated using Pydantic, and generated using the advanced Claude 3.7 Opus model (claude-opus-4-7).
Dataset Structure
The dataset is stored in JSON… See the full description on the dataset page: https://huggingface.co/datasets/himanshunakrani9/principal-systems-architect-dataset.NASA-Systems-Engineering-QA-JSONLsystems-observatory-annotations
Te Pā Systems Observatory · Community Annotations
Public archive of community-submitted rankings of interventions against Donella
Meadows' twelve leverage points, in the Aotearoa New Zealand political-economy
context. Ranked by kaimahi (community members) through the
Systems Observatory dashboard.
Built under the Māori Data Governance Model of
Te Kāhui Raraunga. Public, aggregated, non-personal data only.
Schema (JSON Lines)
Each row in annotations.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/te-pa/systems-observatory-annotations.embedded-systems-qa-dataset4cznz_mechanical_systems_reasoning_eval_v2language:
en
license: other
task_categories:
text-generation
pretty_name: "4CZNZ Mechanical Systems Operational Reasoning Evaluation Sample v2"
size_categories:
n<1K
tags:
operational-cognition
reasoning
operational-reasoning
engineering
troubleshooting
diagnostics
industrial-systems
plc
telemetry
replayability
observability
evaluation-dataset
agentic-systems
rag
autonomous-agents
structured-reasoning
4CZNZ Mechanical Systems Operational Reasoning Evaluation Sample v2
This is a… See the full description on the dataset page: https://huggingface.co/datasets/4CZNZ/4cznz_mechanical_systems_reasoning_eval_v2.control_systems_safety_mathNASA-Systems-Engineering-QA-AlpacaNASA-Systems-Engineering-QA-ChatMLsn96g-coding-systems-2chunk1-20250919_200350
Subnet 96 — Clean Q/A Dataset
Format: one JSONL per line:
{"system": null, "conversations":[{"role":"user","content":"..."}, {"role":"assistant","content":"..."}]}
Total pairs: 2
Avg answer length (tokens): 44 (median 44.0, min 40, max 48)
Schema errors: 0 (should be 0)
File size: 0.00 MB
SHA256 (data.jsonl): 45699c9da9a5509f86d15856917b38f0d4364a1e97e7d30984dffa88158a79ae
Language: English
Intended for: Bittensor Subnet 96 validators
Generation: local LLaMA (GPU) +… See the full description on the dataset page: https://huggingface.co/datasets/raniero/sn96g-coding-systems-2chunk1-20250919_200350.signals-n-systemsspace_systems_qas_datasetThis dataset is obtained from the data for this research paper:https://ieeexplore.ieee.org/stamp/stamp.jsp?arnumber=9548078.
This is an upload of 10K lines of the data with corresponding questions generated for each.
The complete data can be obtained from https://pureportal.strath.ac.uk/en/publications/space-transformers-language-modeling-for-space-systems
NASA-Systems-Engineering-QA-FTSystems-detectivetelugu_kernel_systems_v9sn96g-coding-systems-2chunk1-20250919_182255
Subnet 96 — Clean Q/A Dataset
Format: one JSONL per line:
{"system": null, "conversations":[{"role":"user","content":"..."}, {"role":"assistant","content":"..."}]}
Total pairs: 2
Avg answer length (tokens): 26 (median 26.0, min 13, max 39)
Schema errors: 0 (should be 0)
File size: 0.00 MB
SHA256 (data.jsonl): dc839fd002658139baaf1af4781c5afae16da753251607c667c84b31568ef341
Language: English
Intended for: Bittensor Subnet 96 validators
Generation: local LLaMA (GPU) +… See the full description on the dataset page: https://huggingface.co/datasets/raniero/sn96g-coding-systems-2chunk1-20250919_182255.biblical-systems-methodsn96g-coding-systems-10chunk1-20250919_131658
Subnet 96 — Clean Q/A Dataset
Format: one JSONL per line:
{"system": null, "conversations":[{"role":"user","content":"..."}, {"role":"assistant","content":"..."}]}
Total pairs: 33
Avg answer length (tokens): 117.8 (median 110, min 90, max 164)
Schema errors: 0 (should be 0)
File size: 0.03 MB
SHA256 (data.jsonl): 971a33ac869673ffa81428785853ae1b5c7ec948a22e7c01e60b2deffeda85a7
Language: English
Intended for: Bittensor Subnet 96 validators
Generation: local LLaMA (GPU) +… See the full description on the dataset page: https://huggingface.co/datasets/raniero/sn96g-coding-systems-10chunk1-20250919_131658.sn96g-coding-systems-2chunk1-20250919_191309
Subnet 96 — Clean Q/A Dataset
Format: one JSONL per line:
{"system": null, "conversations":[{"role":"user","content":"..."}, {"role":"assistant","content":"..."}]}
Total pairs: 2
Avg answer length (tokens): 41 (median 41.0, min 40, max 42)
Schema errors: 0 (should be 0)
File size: 0.00 MB
SHA256 (data.jsonl): a1395cd2f98895aba6d95ac3acc1a6852f1332ada930eac8771cd81b9e1121ca
Language: English
Intended for: Bittensor Subnet 96 validators
Generation: local LLaMA (GPU) +… See the full description on the dataset page: https://huggingface.co/datasets/raniero/sn96g-coding-systems-2chunk1-20250919_191309.
