datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
jev-alt-systemone-evalbekko-system-one-v0-decision-index-results
Bekko System One v0 — Decision Index 0.2.1
Evaluation results and reproducibility artifacts for Bekko System One v0
(17M, 68M and 400M) on Decision Index 0.2.1.
This dataset contains predictions, scores, runtime provenance and
training-overlap disclosures for Decision Index evaluation. It does not
redistribute the benchmark inputs.
The runs use the official Decision Index reproduction kit.
These are local strict/no-truncation results, pending maintainer validation
and… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/bekko-system-one-v0-decision-index-results.bekko-system-one-dataset-v0
Bekko System One dataset v0
This dataset is used to train the bekko-system-one-v0 model family.
Public Hugging Face dataset distribution.
Loading and storage
from datasets import load_dataset
dataset = load_dataset("hotchpotch/bekko-system-one-dataset-v0", "laya__typed_decisions")
test = dataset["test"]
Select a configuration explicitly. Parquet uses Zstandard compression and dictionary encoding; all decoded values and row order match the local Arrow release.… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/bekko-system-one-dataset-v0.system-one-decisions
Featuring Labeled Customer Emails and Support Responses
🔧 Synthetic IT Ticket Generator — Custom Dataset
Create a dataset tailored to your own queues & priorities (no PII).
👉 Generate custom data
Define your queues, priorities, language
Need an on-prem AI to auto-classify tickets?→ Open Ticket AI
There are 2 Versions of the dataset, the new version has more tickets, but only languages english and german. So please look at both files, to find what best fits… See the full description on the dataset page: https://huggingface.co/datasets/pngwn/system-one-decisions.systemone-lite-general
systemone-lite-general
Synthetic typed-decision rows for
systemone-lite
(letter-alias choice labels for causal LM SFT).
Not affiliated with TypeSafe AI / Jev. Labels are rule-based, not human prefs.
Splits
Split
Rows
Notes
train
32 400
Stratified mix of 3 gyms
test
3 600
iid held-out by task
test_hard
5 400
layout / paraphrase / option-subset shift
full
36 000
train + iid test
Gyms
TicketDungeon: ticket.route… See the full description on the dataset page: https://huggingface.co/datasets/dwidlee/systemone-lite-general.typed-decisions-v2-system-onesystem-one-270m-data
system-one-270m-data
25,002 synthetic typed decisions: a piece of state, a question, a
caller-supplied option set, and a soft target distribution over those options.
Built to train kaivoss/system-one-270m,
an open take on the System One model class (TypeSafe
Jev,
Laya).
Schema
Field
Type
Meaning
prompt
string
the full rendered prompt, state + question + lettered options
letters
list[string]
the option letters in play, ["A", "B", ...]
target… See the full description on the dataset page: https://huggingface.co/datasets/kaivoss/system-one-270m-data.system-one-mini-data
System One Mini Synthetic Diagnostics
Deterministic synthetic controlled-intervention summaries used by DavidHatley/system-one-mini. The records contain no real coding-agent traces, private repositories, personal data, external model outputs, or teacher labels.
This dataset and model are independent research artifacts, not reproductions of Jev or RLCD.
Splits
Split
Rows
train
20,000
validation
2,000
calibration
2,000
development_renderer
2,000… See the full description on the dataset page: https://huggingface.co/datasets/DavidHatley/system-one-mini-data.systemone-lite-phase2
systemone-lite-phase2
Typed System One distill rows (task / state / instructions / criteria /
label_alias) for systemone-lite.
Revision (2026-09-26)
Paired with Hub model revision action-v2-qwen
(dwidlee/systemone-lite-0.5b).
Change
Detail
Chess
staged_v1 — piece + destination, option caps ≤8
Spatial
action_v2 — legal-only actions; Connect4 drop≤3 + win_now
Sokoban test
Deadlock alerts balanced 75/75 yes/no; remap_alert_prob≈0.35
Postmortem:… See the full description on the dataset page: https://huggingface.co/datasets/dwidlee/systemone-lite-phase2.system-one
Najd System One — public research draft
Decision tasks in English, MSA and Saudi Arabic. This draft is for research and debugging; independent linguistic and label review is pending.
Pack
Cases
License
controlled-development
600
CC-BY-4.0
controlled-validation
600
CC-BY-4.0
controlled-reserved_evaluation
1200
CC-BY-4.0
natural-development
216
CC-BY-4.0
references/arbanking77
1155
CC-BY-SA-4.0
references/massive
576
CC-BY-4.0
references/paired-tool-use
150… See the full description on the dataset page: https://huggingface.co/datasets/najdresearch/system-one.open-system-one-bench
open-system-one-bench — per-item predictions, 10,000 decisions, 5 stacks
Per-example predictions for 10,000 classification/routing decisions, from six
different stacks including typesafe/jev and convaiinnovations/laya.
As far as we know this is the only public per-item output from Jev on a standard
benchmark, and the only place the two are measured on identical items.
With this file you can, without spending a cent on API calls:
run significance tests (McNemar) — is 78.7 % vs… See the full description on the dataset page: https://huggingface.co/datasets/dylantom2012/open-system-one-bench.system-one-training-pairs
System One training pairs
Everything needed to fine-tune the System One cross-encoder on a GPU box, with no git clone.
path
what it is
data/pairs-gold.jsonl
compiled (premise, hypothesis, label) pairs from each dataset's own labels
data/distill/*.jsonl
Claude Haiku 4.5's answers next to gold, one row per training item
data/cache/*.jsonl
the deterministic train/calib/test splits, seed 42
src/
systemone/, tasks/ and train/, enough to run python -m train.finetune… See the full description on the dataset page: https://huggingface.co/datasets/shreyanbr/system-one-training-pairs.SystemOne
SystemOne-4B-Agentic-AGI Dataset
High-quality training & evaluation dataset for a hybrid System One + Agentic model.
This dataset is designed to teach and evaluate:
Fast, calibrated, typed decisions (Choice / Score / Noul)
Parallel multi-question evaluation
Confidence-aware behaviour
Agentic tool use, planning and reflection
Production-style scenarios (support, risk, routing, moderation, etc.)
Design Principles
Atomic & compositional – Prefer narrow… See the full description on the dataset page: https://huggingface.co/datasets/roskosmos19/SystemOne.system-one-eval
System One Eval
Structured State Reasoning Probe, v0.1.
This is a small, openly keyed diagnostic set, not a population benchmark or a blinded leaderboard. It contains 60 originally authored synthetic text tasks and 70 named questions covering policy precedence, cross-row aggregation, joins, boundaries, scheduling, access control, evidence limits, and multi-step arithmetic. There are no images or borrowed public-benchmark items in this package.
What is in the data… See the full description on the dataset page: https://huggingface.co/datasets/blazeofchi/system-one-eval.system-one-datasets
System One Datasets
Typed-decision datasets for System One models, normalized to the /v1/systemone wire format.
Every row is one typed decision (noul, choice, or score) whose state and question, once decoded, are
the body of a POST /v1/systemone request, the API served by TypeSafe's Jev and by open reimplementations such
as openjev. Use the rows for evaluation, calibration, regression tests, or training data selection.
This dataset is not affiliated with or endorsed by TypeSafe… See the full description on the dataset page: https://huggingface.co/datasets/zchee/system-one-datasets.classone-system-one-curriculum
ClassOne System 1 Decision Curriculum (23,503 Examples)
The ClassOne System 1 Decision Curriculum is a multi-domain, structured decision corpus designed to train and benchmark zero-generation System 1 decision models. Instead of generating free-form conversational text, models trained on this curriculum evaluate complex contexts (state) against typed questions (noul, choice, score) in a single forward pass, returning calibrated probabilities, categorical selections, and… See the full description on the dataset page: https://huggingface.co/datasets/devops-thiago/classone-system-one-curriculum.decima-system-one-tasks
Decima System One Tasks
About 138k typed decisions in the Jev / TypeSafe schema, across 15k tasks and nine languages, each with a
calibrated soft label. A task is what a developer writes once: a noul statement to judge true or false, a choice
with named criteria, or a score with ordered levels. Each task comes with many states it is applied to, the way
System One models are used in production.
This is the System One part of the training data of decima-base and
decima-agent. It… See the full description on the dataset page: https://huggingface.co/datasets/amyrmahdy/decima-system-one-tasks.R1-Onevision-with-Systemone-system-knowledge
ONE SYSTEM Knowledge Dataset
Structured knowledge from the ONE SYSTEM ecosystem by David Sanker — a UAPK-centered portfolio of AI, legal, and technology brands.
Training data for LLM research, RAG systems, and AI knowledge graphs. CC-BY-4.0. Updated monthly.
Author
David Sanker — Lawyer (Hucke & Sanker) | AI Engineer (Lawkraft) | Inventor (UAPK)
ORCID: 0009-0004-9636-3910
LinkedIn: linkedin.com/in/sankerlaw
GitHub: github.com/Amakua/one-system-knowledge… See the full description on the dataset page: https://huggingface.co/datasets/LawkraftDavid/one-system-knowledge.vistalab-system-one-synthetic
vistalab-system-one-synthetic
Synthetic Turkish System One questions: a text, a question with named answers, and the probability that open teacher
models give each answer. It is the synthetic half of the training data of
vistalab-system-one-12b and
vistalab-system-one-e4b.
A System One model reads a text and a question and gives a probability to every answer in one forward pass. Questions
are choice (named options, each with a criterion), noul (yes or no) or score (an ordered… See the full description on the dataset page: https://huggingface.co/datasets/UlkuTuncerKucuktas/vistalab-system-one-synthetic.llama2-politosphere-fine-tuning-system-prompt-without-definition
Dataset Card for "llama2-politosphere-fine-tuning-system-prompt_without_definition"
More Information needed
llama2-politosphere-fine-tuning-system-prompt_with_definition
Dataset Card for "llama2-politosphere-fine-tuning-system-prompt_with_definition"
More Information needed
llama2-SST2-SFT-with-system-prompt
Dataset Card for "llama2-SST2-SFT-with-system-prompt"
More Information needed
llama2-politosphere-fine-tuning-system-prompt
Dataset Card for "llama2-politosphere-fine-tuning-system-prompt"
More Information needed
llama2-sst2-fine-tuning-without-system-info
Dataset Card for "llama2-sst2-fine-tuning-without-system_info"
More Information needed
