Team Ai
Datasetpublic

muahmed7338/kernelascent-tasks

KernelAscent — public dev split KernelAscent is a benchmark for recursive self-improvement (RSI): a model optimizes the GPU kernels used to train itself, and we measure whether kernel-optimization capability compounds across rounds. This is the public dev split, released for self-benchmarking and research; the leaderboard is scored on a private held-out split. Project & code: https://github.com/ahmd-mohsin/KernelAscent Leaderboard & docs:… See the full description on the dataset page: https://huggingface.co/datasets/muahmed7338/kernelascent-tasks.

sourceHugging Facemitupdated 22d agoView on Hugging Face
0likes548downloads
Dataset Card

KernelAscent — public dev split

KernelAscent is a benchmark for recursive self-improvement (RSI): a model optimizes the GPU kernels used to train itself, and we measure whether kernel-optimization capability compounds across rounds. This is the public dev split, released for self-benchmarking and research; the leaderboard is scored on a private held-out split.

  • —Project & code: https://github.com/ahmd-mohsin/KernelAscent
  • —Leaderboard & docs: https://ahmd-mohsin.github.io/KernelAscent/

Internal-failure causality DAG

Mechanistic interpretability of RSI: every model traces one path through the internal gates it must clear — scale → correctness-wall → gradient → drift → retention → diversity → outcome (green = RSI compounds, blue = crossed the wall but flat, orange = stuck at the wall). Sub-2B models stall at the correctness wall; mid-scale (2–8B) threads every gate and compounds; the largest drift most yet saturate at the roofline.

Scale vs RSI gain with marginals

Held-out capability gain vs model size, bubble area ∝ LoRA drift. Shaded = sub-2B correctness wall; positive gain concentrates at mid-scale. Interactive versions on the [project site](https://ahmd-mohsin.github.io/KernelAscent/).

Key findings (updated 2026-09-18)

*Headline: self-training sharpens what a model already covers; it does not explore. In GPU-kernel optimization, compounding is coverage-limited, and coverage is not self-generated.*

  • —No compounding (strong, well-powered null). Lineage vs.\ matched-reset over multi-round rejection-sampling SFT: pooled lineage$-$reset $=+0.001\,[-0.015,+0.018]$ ($n{=}213$), TOST-equivalent at $\delta{=}0.05$, Bayes factor $\mathrm{BF}_{01}\approx14.5$ (strong evidence for the null); every scale 0.5–3B individually equivalent. 7B/14B extension in progress.
  • —Search beats training (significant). At matched compute, best-of-$N$ search beats lineage self-training: lineage$-$bestof$N=-0.113\,[-0.153,-0.074]$ (CI excludes 0); search wins 70% of rounds.
  • —Coverage-vs-scale curve (0.5B→14B). Coverage 14%→86%; below the wall pass@$K\gg$pass@1 (12–17×) — the sub-2B "correctness wall" is a sampling artifact, not absent capability.
  • —0-score failure forensics (8 model families). The sub-3B wall is a kernel-formation failure (no_extract 57–96%: the model never emits a valid kernel), not a correctness failure; the failure locus marches downstream with scale (incoherence → truncation → API-hallucination → wrong-output → correct).
  • —Weight-level mechanism (WHY-RSI, $n{=}133$). RSI is an inverted-U in scale: peaks at 2–8B (49%), while ≥9B shows the lowest RSI (19%) despite the highest LoRA drift (0.73) — large models churn weights without compounding (roofline saturation). Drift localizes to late layers.
  • —Task-5 self-play (closed-source frontier models). Self-modification gain is one-shot (68% at round 0), 75% of models non-recursive; one model self-degrades.

Raw per-run trajectories (compounding histories, WHY-RSI series, forensics reasoning-chains, closed-source self-modify) are in `data/trajectories/` for independent analysis.

What is in a task

Each task is a self-contained, seeded PyTorch Model whose forward is a fused op-graph; an agent must return an optimized, numerically-equivalent ModelNew (Triton or fused PyTorch). Per-task files:

  • —task.py — the problem (Model, get_inputs, seeded weights).
  • —meta.json — tier, family, tags, shape/dtype/chain, and (curation) achievable_speedup, pass_rate, difficulty.
  • —reference_solution.py — the best correct + fastest kernel found by the curator (Claude Fable 5). The achievable target.
  • —results.json — full grading record (per-candidate correctness, timing, speedup vs eager and vs the min(eager, torch.compile) roofline).

Structure: difficulty tiers with empirical labels

The public split is organized by difficulty tier under public/<Tier>/<task>/:

  • —Easy: small power-of-two elementwise fusion or a single reduction (softmax, layernorm, rmsnorm). Accessible floor.
  • —Medium: matmul with a fused epilogue, or short fused chains.
  • —Hard: matmul-bearing chains, full and causal attention, RoPE attention.
  • —Ultra: soft-MoE and large or irregular shapes.

Every task's meta.json carries an empirical difficulty measured by running 13 open-weight models (Qwen2.5-Coder / Qwen2.5-Instruct 0.5B to 14B, DeepSeek-Coder-6.7B, StarCoder2-15B, CodeLlama-13B): solve_rate (fraction of models that produced a correct kernel) and best_speedup_observed (best speedup vs the min(eager, torch.compile) roofline any model achieved), plus a difficulty label (speed-open, correctness-only, hard, unsolved). public/manifest.json indexes the whole set. Empirical difficulty distribution:

Easy    25 speed-open, 5 correctness-only
Medium  18 speed-open, 10 correctness-only, 2 rare
Hard    11 speed-open, 18 correctness-only, 1 hard
Ultra    8 speed-open, 16 correctness-only, 4 hard, 2 rare

Correctness difficulty rises monotonically Easy to Ultra. The roofline is torch.compile, so there is real headroom above the bar at every tier (no global optimum). See the repo analysis/calibration_run.md for the failure breakdown.

How we evaluate

Correctness. A candidate ModelNew is checked against an fp32 gold on N=4 fresh random inputs with a dtype-aware tolerance and an input-sensitivity check that rejects constant or input-ignoring outputs. Correctness is verified on the timed run. Each candidate is graded in an isolated subprocess so a native compiler abort or hang loses only that candidate.

Two walls, reported separately. Correctness rate (was a valid correct kernel produced) and speed rate (does a correct kernel beat the roofline). We never fuse them into one number.

Speed score. Continuous log-interpolated ladder between eager, torch.compile, and an expert kernel: s = clip((ln t_eager - ln t_cand)/(ln t_eager - ln t_expert), 0, 1.2), 0 at eager, 1 at expert, compile parity as a milestone. Expert rungs are reconstructed with a strong curator (Fable 5.1) and verified to beat torch.compile.

How progress (RSI) is measured

Capability is the tier ladder. Recursive self-improvement is measured causally. A 15-model x 4-arm sweep (growing / frozen-nonempty / offline-built / matched-search) found matched-compute search beats recursive library-growing on average (growing below its strongest control for 13 of 15 models): a clean negative for memory-RSI on this benchmark. The v3 redesign makes the central object the causal returns to recursive improvement: separate the actor (the procedure producing a patch) from the target (what is patched) so competing producers edit the SAME target, and measure Q (research productivity), V (producing a better improver), and the causal producer contrast F across a two-link lineage with rescue. A deterministic calibration suite proves the instrument distinguishes a repeating recursive positive control from a one-time upgrade, best-of-N, and nulls before any model is judged. Full design in the project repo docs/RSI_V3_PLAN.md. The private held-out split is not released.

Families (6) and tiers

matmul (L2), norm-act (L1), attention (L3), rope-attention (L3), quant-gemm (L2, int8 dequant + GEMM), moe (L3, gated experts / grouped GEMM). Tiers: L1 memory/reduction, L2 tensor-core/matmul-epilogue, L3 attention & structured.

Scoring

  • —Correctness against an fp32 gold, allowing no more error than the working fp16/bf16 dtype itself incurs.
  • —Roofline-relative speedup t_baseline / t_candidate, baseline = min(eager, torch.compile).
  • —fast_p (fraction beating p× speedup) and pass@k; timing is warmup + median-of-N + L2 flush on clock-pinned GPUs.

Provenance & contamination

Tasks are synthesized deterministically from seeds at generation time (not drawn from a fixed public list). Public and private held-out seed ranges are disjoint; the held-out split is never released, so leaderboard scores cannot be gamed by overfitting the public set.

Citation

@misc{kernelascent2026,
  title  = {KernelAscent: Measuring Recursive Self-Improvement via a Kernel-to-Model Capability Loop},
  author = {Mohsin, Ahmed},
  year   = {2026},
  url    = {https://github.com/ahmd-mohsin/KernelAscent}
}