callensxavier/laya-coding-curriculum-78k
Laya-Curriculum-78k: 10 Structured Coding Datasets + 5 Verifier Datasets Unified, multi-task curriculum dataset designed for training non-autoregressive System 1 code companions, quality gates, and neuro-symbolic routing models (Laya-LoRA and GWAYA). Total Core Records: 78,503 structured examples across 3 stages Verifier Records: 500 multi-domain verifier examples (Python, Rust, Lean 4, Mathematics) Target Model: callensxavier/laya-lora-modernbert-r8 Permanent Archival DOI:… See the full description on the dataset page: https://huggingface.co/datasets/callensxavier/laya-coding-curriculum-78k.
Laya-Curriculum-78k: 10 Structured Coding Datasets + 5 Verifier Datasets
Unified, multi-task curriculum dataset designed for training non-autoregressive System 1 code companions, quality gates, and neuro-symbolic routing models (Laya-LoRA and GWAYA).
- Total Core Records: 78,503 structured examples across 3 stages
- Verifier Records: 500 multi-domain verifier examples (Python, Rust, Lean 4, Mathematics)
- Target Model: callensxavier/laya-lora-modernbert-r8
- Permanent Archival DOI: 10.5281/zenodo.23076256
- License: Apache 2.0
1. Curriculum Stages and Subsets
Stage 1: Security, Vulnerability, and Anti-Stub (46,196 records)
Focuses on binary gating ($p_\text{pass} \in [0, 1]$), blocking malicious patterns, stub detection (pass, ..., unproven sorry, unimplemented!()), and security classification.
- PyCode-Vul (28,340 records): Real-world Python vulnerabilities and exploits.
- CodeRM-UnitTest (17,562 records): Unit tests with execution assertions.
- SmellBench (294 records): Code smells, God classes, and architectural anti-patterns.
Stage 2: Performance Efficiency & Energy (1,263 records)
Focuses on predicting execution complexity ($O(1)$ to $O(n^3)$) and continuous physical energy regression ($E \in [0, \infty)$ Joules).
- EffiBench-X (623 records): Algorithmic execution efficiency.
- SWE-Perf (140 records): Real-world software engineering performance patches.
- RAPL Energy Bench (500 records): Hardware energy consumption measurements.
Stage 3: Formal Verification & Trace Alignment (31,044 records)
Focuses on formal mathematical proofs, interactive theorem proving, and execution alignment.
- Lean-Workbook (10,000 records): Formal Lean 4 mathematical proofs.
- miniF2F-Lean4 (244 records): Formal olympiad math benchmark problems.
- Magpie-Qwen2.5-20K (20,000 records): High-quality multi-turn reasoning traces.
- CRUXEval (800 records): Code reasoning and input/output execution prediction.
Verifier Curriculum: GWAYA Multi-Domain Ground Truth (500 records)
Targeted verification records for parallel zero-trust verification of frontier agent outputs (Gemini 3.1 Pro & 3.8 Flash):
- Muennighoff/mbpp (Python: verified unit tests & assertions)
- bigcode/the-stack-smol-rust (Rust: borrow checking & SIMD memory safety)
- HuggingFaceH4/MATH-500 (Formal math & Lean 4 olympiad reasoning)
- semeru/code-smell-dataset (Python: AST smells & stub detection)
- openai/gsm8k (Step-by-step arithmetic & numerical consistency)
2. Record Schema
Each record is formatted as a single JSON line:
{
"text": "def compute(x): return x * 2",
"noul_label": 1,
"choice_label": 0,
"score_label": 0.05,
"gate_score": 1.0,
"dataset_id": "clean_code",
"pillar": "python"
}3. Citation
@article{callens2026laya,
title={Laya-LoRA Coding Companion: Asymmetric Dual-Process Test-Time Compute, Full-Scale Curriculum on 10 Structured Coding Datasets, and Serverless Multi-Tier Inference Infrastructure on GCP},
author={Callens, Xavier},
journal={Systems for Machine Learning / Compound AI Systems},
year={2026},
doi={10.5281/zenodo.23076256},
url={https://huggingface.co/datasets/callensxavier/laya-coding-curriculum-78k}
}