datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kernelbench-mega-traces
KernelBench-Mega agent traces
Coding agents writing full GPU megakernels across Blackwell / H100 / B200, scored as speedup over reference; contamination-audited (23 verified cells).
Each .jsonl file is one agent run in Claude-Code session format, viewable with the agent trace viewer. Filename = run id; manifest.csv maps each run to model / harness / problem / GPU / score.
23 agent traces · live leaderboard: https://kernelbench.com/mega
Secrets redacted. Full reasoning for… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-mega-traces.kernelbench-hard-traces
KernelBench-Hard agent traces
Frontier coding agents writing optimized CUDA/Triton kernels (FP8 GEMM, paged
attention, MoE, W4A16, KDA, Top-k) on RTX PRO 6000 Blackwell, H100 PCIe, and
B200; roofline-graded.
Each .jsonl file is one agent run in Claude-Code session format, viewable with
the Hugging Face Agent Trace viewer (Data Studio → open a row). Filename =
run id.
Live leaderboard: https://kernelbench.com/hard
Secrets redacted. Full reasoning for open-provider routes… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-hard-traces.kernelbench-cuda-traceskernelbench-hard-runs
KernelBench-Hard — Agent Runs
84 full agent transcripts (12 frontier models × 7 problems) from the KernelBench-Hard sweep on a single Blackwell GPU (RTX PRO 6000, sm_120, CUDA 13.2). Each run contains the model's full reasoning trace, every tool call, the final solution.py, and the eval result.
Companion datasets:
Infatoshi/kernelbench-hard-problems — the 7 problem definitions
Live site: https://kernelbench.com/hard
100 themed transcript viewers (HTML): https://kernelbench.com/runs… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-hard-runs.gspc-kernel-results
Kaggle 3090 ladder — prompt-bank passes
Prompt-bank passes from the Kaggle 3090 ladder. Each row of
kernel_results.jsonl carries the axis, the model family and full model id, the prompt, the score,
the n behind that score, a sigil content hash, the platform it ran on and the timestamp. Most rows are
n=1 single-prompt passes — read them as a ladder sweep across many open models, not as board n.
The live board is the authority
GET https://councilof.ai/api/gspc —… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-kernel-results.kern-kernels
kern-kernels
Reproducible attention kernel recipes, ABI manifests, checksums and measured
results. First profile: GB300 / Qwen3.8-27B / BF16 / Q24-KV4-D256 / page 64.
Uses unmodified TRTLLM-GEN full attention, not MLA. The model's GDN layers are
unchanged. Model weights and the base export's other kernels are not included.
NVIDIA binaries are downloaded directly from pinned upstream URLs and verified
by SHA256; this repository does not mirror them. The small Apache-2.0 vLLM KV… See the full description on the dataset page: https://huggingface.co/datasets/susun-123/kern-kernels.glm-5.2-kernelgym-rollouts
GLM-5.2 KernelGym Rollouts
This dataset contains 3,200 feedback-driven GPU-kernel optimization trajectories
generated by zai-org/GLM-5.2-FP8: 100 validation tasks, two backends (inline
CUDA and Triton), and 16 rollouts per task.
Each trajectory retains the prompt/feedback message history, model responses and
reasoning, extracted kernel code, KernelGym compilation and correctness results,
profiling metadata, token usage, and stopping decision. Every published record
ended with… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/glm-5.2-kernelgym-rollouts.kernelbench_with_promptsThis is a version of KernelBench where the prompts to produce the Triton and cuda kernel are explicitly saved in the JSON data files.
It only contains Level 1, 2, 3 kernels.
The prompt is the same as what is provided in the original KernelBench repo.
The dataset is prepared by Jiin Woo during her internship at AWS Annapurna Labs, the lab behind Trainium chips.
This dataset is part of an unreleased paper, and the paper will be updated in this README soon. If you use this dataset, please cite… See the full description on the dataset page: https://huggingface.co/datasets/allenanie/kernelbench_with_prompts.kernel-grader-mi300x-rocm-trainium-nki
Grading a kernel on an MI300X, and two agents that skipped it entirely
I put 67,108,864 skewed keys into 4096 bins on an AMD Instinct MI300X VF, gfx942, under ROCm 7.2.4. One global atomic per key takes 86.3898 ms. Privatising the table into local data share per block, with one merge at the end, takes 0.1015 ms. That is 851.13 times, measured on the same card in the same run.
Two cheating agents then scored full marks against my own grader.
What ran
Measured time
Score… See the full description on the dataset page: https://huggingface.co/datasets/LaelaZorana/kernel-grader-mi300x-rocm-trainium-nki.litmus-kernels
Litmus Kernel Verification Corpus
Correct and deliberately-broken Triton kernels, each broken one shipped with
the input that exposes it.
The corpus exists to measure one thing: how much of what a fixed-shape
torch.rand() allclose test calls "correct" actually is. On this corpus the
answer is that 88% of the planted bugs pass that test.
Columns
column
meaning
name
kernel identifier
family
elementwise / reduction / softmax / layernorm / matmul /… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/litmus-kernels.b-sides-v2-kernels
B-Sides kernel search index
Published runtime artifacts for the B-Sides kernel corpus.
meta.sqlite: searchable kernel metadata
embeddings.f16.npy: row-aligned 256-dimensional query matrix
The kernels.row values align one-to-one with the embedding matrix rows.
tpu-kernel-parity-lab-v5e
JAX/Pallas attention measured on one TPU v5e device
I measured a Pallas masked-softmax kernel against the JAX/XLA reference inside grouped-query
attention. The Kaggle worker exposed eight TPU devices. The unsharded arrays ran on JAX's default
single device, so these numbers make no multi-device scaling claim.
All seven correctness checks passed. The benchmark contains twenty cases, with five warmups and
twenty synchronized timing samples in each case. XLA was faster in all ten… See the full description on the dataset page: https://huggingface.co/datasets/LaelaZorana/tpu-kernel-parity-lab-v5e.kernelx-training-dataaetherbrowser-kernel
AetherBrowser kernel evidence packet
This public dataset is a content-addressed, secret-free evidence mirror for the
AetherBrowser transaction kernel. It is not a training dataset, hosted inference
service, credential store, or competition submission.
The kernel constrains browser work to:
observe -> plan -> approve -> dispatch -> verify -> receipt
Raw page text, fill values, and credentials are excluded from durable state. Unknown
operations and remote-write flows fail closed.… See the full description on the dataset page: https://huggingface.co/datasets/issdandavis/aetherbrowser-kernel.rope-kernel-4b-abrope-kernel-v2Kernel-Smith-Seed-59KIf this work is useful to you, please cite:
@article{DBLP:journals/corr/abs-2603-28342,
author = {He Du and
Qiming Ge and
Jiakai Hu and
Aijun Yang and
Zheng Cai and
Zixian Huang and
Sheng Yuan and
Qinxiu Cheng and
Xinchen Xie and
Yicheng Chen and
Yining Li and
Jiaxing Xie and… See the full description on the dataset page: https://huggingface.co/datasets/CoopReason/Kernel-Smith-Seed-59K.test_kernelbookkernelbook-test
