Team Ai
Datasetpublic

callensxavier/laya-coding-curriculum-78k

Laya-Curriculum-78k: 10 Structured Coding Datasets + 5 Verifier Datasets Unified, multi-task curriculum dataset designed for training non-autoregressive System 1 code companions, quality gates, and neuro-symbolic routing models (Laya-LoRA and GWAYA). Total Core Records: 78,503 structured examples across 3 stages Verifier Records: 500 multi-domain verifier examples (Python, Rust, Lean 4, Mathematics) Target Model: callensxavier/laya-lora-modernbert-r8 Permanent Archival DOI:… See the full description on the dataset page: https://huggingface.co/datasets/callensxavier/laya-coding-curriculum-78k.

sourceHugging Faceapache-2.0updated 9d agoView on Hugging Face
0likes66downloads
Dataset Card

Laya-Curriculum-78k: 10 Structured Coding Datasets + 5 Verifier Datasets

Unified, multi-task curriculum dataset designed for training non-autoregressive System 1 code companions, quality gates, and neuro-symbolic routing models (Laya-LoRA and GWAYA).


1. Curriculum Stages and Subsets

Stage 1: Security, Vulnerability, and Anti-Stub (46,196 records)

Focuses on binary gating ($p_\text{pass} \in [0, 1]$), blocking malicious patterns, stub detection (pass, ..., unproven sorry, unimplemented!()), and security classification.

  • —PyCode-Vul (28,340 records): Real-world Python vulnerabilities and exploits.
  • —CodeRM-UnitTest (17,562 records): Unit tests with execution assertions.
  • —SmellBench (294 records): Code smells, God classes, and architectural anti-patterns.

Stage 2: Performance Efficiency & Energy (1,263 records)

Focuses on predicting execution complexity ($O(1)$ to $O(n^3)$) and continuous physical energy regression ($E \in [0, \infty)$ Joules).

  • —EffiBench-X (623 records): Algorithmic execution efficiency.
  • —SWE-Perf (140 records): Real-world software engineering performance patches.
  • —RAPL Energy Bench (500 records): Hardware energy consumption measurements.

Stage 3: Formal Verification & Trace Alignment (31,044 records)

Focuses on formal mathematical proofs, interactive theorem proving, and execution alignment.

  • —Lean-Workbook (10,000 records): Formal Lean 4 mathematical proofs.
  • —miniF2F-Lean4 (244 records): Formal olympiad math benchmark problems.
  • —Magpie-Qwen2.5-20K (20,000 records): High-quality multi-turn reasoning traces.
  • —CRUXEval (800 records): Code reasoning and input/output execution prediction.

Verifier Curriculum: GWAYA Multi-Domain Ground Truth (500 records)

Targeted verification records for parallel zero-trust verification of frontier agent outputs (Gemini 3.1 Pro & 3.8 Flash):

  1. 1.Muennighoff/mbpp (Python: verified unit tests & assertions)
  2. 2.bigcode/the-stack-smol-rust (Rust: borrow checking & SIMD memory safety)
  3. 3.HuggingFaceH4/MATH-500 (Formal math & Lean 4 olympiad reasoning)
  4. 4.semeru/code-smell-dataset (Python: AST smells & stub detection)
  5. 5.openai/gsm8k (Step-by-step arithmetic & numerical consistency)

2. Record Schema

Each record is formatted as a single JSON line:

json
{
  "text": "def compute(x): return x * 2",
  "noul_label": 1,
  "choice_label": 0,
  "score_label": 0.05,
  "gate_score": 1.0,
  "dataset_id": "clean_code",
  "pillar": "python"
}

3. Citation

bibtex
@article{callens2026laya,
  title={Laya-LoRA Coding Companion: Asymmetric Dual-Process Test-Time Compute, Full-Scale Curriculum on 10 Structured Coding Datasets, and Serverless Multi-Tier Inference Infrastructure on GCP},
  author={Callens, Xavier},
  journal={Systems for Machine Learning / Compound AI Systems},
  year={2026},
  doi={10.5281/zenodo.23076256},
  url={https://huggingface.co/datasets/callensxavier/laya-coding-curriculum-78k}
}