datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
scbe-aethermoore-training-data
Status: canonical. Primary public training dataset for SCBE-AETHERMOORE and the most-used repo in this account. Other scbe-* dataset repos are experiment-specific slices.
SCBE-AETHERMOORE Training Dataset
Supervised fine-tuning (SFT) dataset for the SCBE-AETHERMOORE hyperbolic geometry AI safety and governance framework.
Overview
This dataset contains 10,978 training pairs spanning the full SCBE-AETHERMOORE system: 14-layer architecture knowledge, Six Sacred… See the full description on the dataset page: https://huggingface.co/datasets/issdandavis/scbe-aethermoore-training-data.AetherCode
AetherCode: Evaluating LLMs' Ability to Win In Premier Programming Competitions
Introduction
Competitive programming has emerged as a critical benchmark for evaluating the reasoning and coding capabilities of Large Language Models (LLMs). Despite impressive progress on existing benchmarks, we argue that current evaluations overstate model proficiency, masking a substantial gap between LLMs and elite human programmers. This gap arises… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/AetherCode.akie-pretrain-corpus
AKIE Pretrain Corpus
Corpus de pré-treino para a família de modelos AKIE, organizado em 4
eixos com proporções fixas, tokenizado com o
AkieTokenizer
(SentencePiece BPE, 32k vocabulário).
Composição
Eixo
Proporção
Tokens
Código
40%
~2,40B
Instruções
20%
~1,20B
Diálogo
25%
~1,50B
Raciocínio
15%
~0,90B
Total
100%
~6,00B
Fontes: código-fonte de repositórios públicos (várias linguagens),
diálogos e instruções em português (traduções e coleções… See the full description on the dataset page: https://huggingface.co/datasets/AETHER-LAB/akie-pretrain-corpus.cpp_cwe_GRPO
cpp_cwe_GRPO
VeRL/GRPO-ready C++ security-coding dataset produced from every harness-passing stage-6 rewrite in the simple_gen pipeline.
Files
File
Rows
Description
cpp_cwe_GRPO.parquet
8,039
Full corrected dataset
cpp_cwe_GRPO_train.parquet
7,236
Deterministic 90% training split
cpp_cwe_GRPO_val.parquet
803
Deterministic 10% validation split
The split uses seed 20260929. Each record has a unique function name, and the two splits are disjoint.… See the full description on the dataset page: https://huggingface.co/datasets/AetherPrior/cpp_cwe_GRPO.c_cwe_GRPO
c_cwe_GRPO
VeRL/GRPO-ready C security coding dataset generated by the simple_gen pipeline.
Each row is a harness-validated task with pytest security/functionality tests and oracle candidate_c. Most CWEs are post stage-6 rubric rewrite (6_rewrites.jsonl); CWE-476 is still from stage-5 guidelines. Stage-7 guideline resampling has not been applied yet, so high_level_guidelines / implementational may be empty on rewritten rows.
Files
File
Rows
Description… See the full description on the dataset page: https://huggingface.co/datasets/AetherPrior/c_cwe_GRPO.AetherCode-v1
Dataset Description
Abstract
The "AetherCode" dataset is designed to fine-tune models on coding tasks across various programming languages, incorporating complex real-world coding scenarios. It aims to push the boundaries of AI in code generation and software development.
How to Load This Dataset
from datasets import load_dataset
dataset = load_dataset("thesven/AetherCode-v1", split="5star")
Languages
The dataset includes coding problems in… See the full description on the dataset page: https://huggingface.co/datasets/thesven/AetherCode-v1.AetherSearch_DPO
🔭 AetherSearch DPO
Preference pairs for reasoning, retrieval, and evidence-grounded answers
🏠 Project ·
🧠 DPO Model ·
🧪 Training Code ·
🎓 SFT Data ·
🤖 SFT Model
Dataset overview
AetherSearch DPO contains 2,126 preference pairs for training an agentic-search
policy after supervised fine-tuning. Every row provides one shared prompt, a
preferred assistant continuation, and a non-preferred continuation.
This repository is… See the full description on the dataset page: https://huggingface.co/datasets/muradil211/AetherSearch_DPO.aether-cyber-sft
enosislabs/aether-cyber-sft
Version: surface-clean-20260620
Generated UTC: 2026-06-20T03:36:32.011333+00:00
Source path: artifacts/aether-cyber-sft-surface-candidate.jsonl
Git commit: 6fadc0ba97b4c49a7b644e70afacddc8ba98029e
Curated Aether PRISM SFT dataset for authorized cybersecurity training.
Each source shard follows 5-row review discipline before publish.
Records
Total examples: 1794
Domains
vulnerability_research: 757
red_team_ops: 334… See the full description on the dataset page: https://huggingface.co/datasets/enosislabs/aether-cyber-sft.kotodama-3b-corpus-open
kotodama-3b pretraining corpus — open subset
The deduplicated, cleaned text of the openly licensed sources used to pretrain the kotodama 3B
(aethera-gp/kotodama-3b-base-final). 19 of the model's 32 sources; the other 13 (non-commercial,
unlicensed, gated, or flagged sources) are not redistributed here.
Layout: data/<source>/part-NNNN.jsonl.zst — zstd JSON Lines, one document per line.
Processing (github.com/LuxiaSL/kotodama, curation): Unicode normalisation, length filtering… See the full description on the dataset page: https://huggingface.co/datasets/aethera-gp/kotodama-3b-corpus-open.aether-family-trading-shared
aether-family-trading-shared
AETHER family SFT dataset — group trading.
Format: JSONL ChatML messages, task_type tagged, MinHash dedup applied (threshold 0.85).
Schema:
{
"messages": [{"role": "system|user|assistant", "content": "..."}],
"task_type": "function_calling|code|reasoning_cot|...",
"source_ds": "<HF dataset_id>",
"lang": "en|fr|...",
"system_source": "archon_default|overridden_from_source"
}
Generated by prepare_sft.py pipeline (2026-05-25).
aether-build-protocol-examples
Aether Build Protocol Examples
Aether Build Protocol Examples is a small public dataset of machine-readable physical build intent artifacts.
It is designed for AI developers, agent-framework builders, CAD/design workflows, fabrication review systems, and researchers studying machine-to-machine physical transaction protocols.
GitHub source of truth:
https://github.com/chevy155/Aether-build-protocol
Live demo:
https://huggingface.co/spaces/lonestar155/aether-cad-to-agent-sandbox
Open… See the full description on the dataset page: https://huggingface.co/datasets/lonestar155/aether-build-protocol-examples.aetherstory-data
AetherStory Dataset
A procedurally generated corpus of unique fantasy fables, purpose-built to
train the AetherStory storyteller
model. Every story is synthesised on the fly from a combinatorial space of
realms, creatures, character archetypes, magic systems, and plot scaffolds.
Why this dataset is unique
It is not scraped from the web. Each fable is constructed by code from
hand-written ingredients, so:
there is no copyright risk — every word is original or… See the full description on the dataset page: https://huggingface.co/datasets/wincode/aetherstory-data.aether-sft-v2-mix
aether-sft-v2-mix
AETHER family SFT dataset — group mix.
Format: JSONL ChatML messages, task_type tagged, MinHash dedup applied (threshold 0.85).
Schema:
{
"messages": [{"role": "system|user|assistant", "content": "..."}],
"task_type": "function_calling|code|reasoning_cot|...",
"source_ds": "<HF dataset_id>",
"lang": "en|fr|...",
"system_source": "archon_default|overridden_from_source"
}
Generated by prepare_sft.py pipeline (2026-05-25).
aether-sft-v2-code
aether-sft-v2-code
AETHER family SFT dataset — group code.
Format: JSONL ChatML messages, task_type tagged, MinHash dedup applied (threshold 0.85).
Schema:
{
"messages": [{"role": "system|user|assistant", "content": "..."}],
"task_type": "function_calling|code|reasoning_cot|...",
"source_ds": "<HF dataset_id>",
"lang": "en|fr|...",
"system_source": "archon_default|overridden_from_source"
}
Generated by prepare_sft.py pipeline (2026-05-25).
js_cwe_GRPO
js_cwe_GRPO
VeRL/GRPO-ready JavaScript security coding dataset generated by the simple_gen pipeline.
Each row is a harness-validated task with Node harness security/functionality tests, oracle candidate_js, and authoring guidelines (high_level_guidelines, implementational).
Files
File
Rows
Description
js_cwe_GRPO.parquet
4956
Full dataset (shuffled)
js_cwe_GRPO_train.parquet
4461
90% train split
js_cwe_GRPO_val.parquet
495
10% validation split… See the full description on the dataset page: https://huggingface.co/datasets/AetherPrior/js_cwe_GRPO.aether-0.8b-cyber-sft
enosislabs/aether-0.8b-cyber-sft
Version: aether-0.8b-cyber-20260618
Generated UTC: 2026-06-19T02:53:44.849450+00:00
Source path: data/curated
Git commit: 3ac5f50de88e43122cfda43d65f1ad01b5febff6
Curated Aether PRISM SFT dataset for authorized cybersecurity training.
Each source shard follows 5-row review discipline before publish.
Records
Total examples: 1900
Domains
vulnerability_research: 715
red_team_ops: 335
detection_engineering: 257… See the full description on the dataset page: https://huggingface.co/datasets/enosislabs/aether-0.8b-cyber-sft.aether-sft-v2-trading
aether-sft-v2-trading
AETHER family SFT dataset — group trading.
Format: JSONL ChatML messages, task_type tagged, MinHash dedup applied (threshold 0.85).
Schema:
{
"messages": [{"role": "system|user|assistant", "content": "..."}],
"task_type": "function_calling|code|reasoning_cot|...",
"source_ds": "<HF dataset_id>",
"lang": "en|fr|...",
"system_source": "archon_default|overridden_from_source"
}
Generated by prepare_sft.py pipeline (2026-05-25).
aether-sft-v2-multilingual
aether-sft-v2-multilingual
Multilingual EN+FR subset from CohereForAI/aya_dataset.
Total: 5366 samples ChatML.
aether-sft-v2-reasoning
aether-sft-v2-reasoning
AETHER family SFT dataset — group reasoning.
Format: JSONL ChatML messages, task_type tagged, MinHash dedup applied (threshold 0.85).
Schema:
{
"messages": [{"role": "system|user|assistant", "content": "..."}],
"task_type": "function_calling|code|reasoning_cot|...",
"source_ds": "<HF dataset_id>",
"lang": "en|fr|...",
"system_source": "archon_default|overridden_from_source"
}
Generated by prepare_sft.py pipeline (2026-05-25).
