datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kernelbench-mega-traces
KernelBench-Mega agent traces
Coding agents writing full GPU megakernels across Blackwell / H100 / B200, scored as speedup over reference; contamination-audited (23 verified cells).
Each .jsonl file is one agent run in Claude-Code session format, viewable with the agent trace viewer. Filename = run id; manifest.csv maps each run to model / harness / problem / GPU / score.
23 agent traces · live leaderboard: https://kernelbench.com/mega
Secrets redacted. Full reasoning for… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-mega-traces.kernelbench-hard-traces
KernelBench-Hard agent traces
Frontier coding agents writing optimized CUDA/Triton kernels (FP8 GEMM, paged
attention, MoE, W4A16, KDA, Top-k) on RTX PRO 6000 Blackwell, H100 PCIe, and
B200; roofline-graded.
Each .jsonl file is one agent run in Claude-Code session format, viewable with
the Hugging Face Agent Trace viewer (Data Studio → open a row). Filename =
run id.
Live leaderboard: https://kernelbench.com/hard
Secrets redacted. Full reasoning for open-provider routes… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-hard-traces.kernelbench-v3-runs
KernelBench-v3 — Agent Runs
2071 agent evaluations from the v3 sweep (2026-02): 10 frontier models × {RTX 3090, H100, B200} × 43–58 problems per GPU. Each row is one (model, gpu, problem) triple with correctness, speedup, baseline timing, token usage, cost, and a pointer to the agent's winning solution.py.
Companion datasets:
Infatoshi/kernelbench-v3-problems — 60 problem definitions
Infatoshi/kernelbench-hard-runs — newer KernelBench-Hard sweep (12 models × 7 problems on Blackwell… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-v3-runs.kernelbench-cuda-tracesKernelBench
KernelBench
A benchmark designed to evaluate the ability of LLMs to generate efficient GPU kernels for optimizing neural network performance
Version
[07-21-2025] This HF dataset version has been updated to v0.1
Citation
@misc{ouyang2024kernelbench,
title={KernelBench: Can LLMs Write GPU Kernels?},
author={Anne Ouyang and Simon Guo and Azalia Mirhoseini},
year={2024},
url={https://scalingintelligence.stanford.edu/blogs/kernelbench/},
}
kernelbench-hard-runs
KernelBench-Hard — Agent Runs
84 full agent transcripts (12 frontier models × 7 problems) from the KernelBench-Hard sweep on a single Blackwell GPU (RTX PRO 6000, sm_120, CUDA 13.2). Each run contains the model's full reasoning trace, every tool call, the final solution.py, and the eval result.
Companion datasets:
Infatoshi/kernelbench-hard-problems — the 7 problem definitions
Live site: https://kernelbench.com/hard
100 themed transcript viewers (HTML): https://kernelbench.com/runs… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-hard-runs.kernelbench_with_promptsThis is a version of KernelBench where the prompts to produce the Triton and cuda kernel are explicitly saved in the JSON data files.
It only contains Level 1, 2, 3 kernels.
The prompt is the same as what is provided in the original KernelBench repo.
The dataset is prepared by Jiin Woo during her internship at AWS Annapurna Labs, the lab behind Trainium chips.
This dataset is part of an unreleased paper, and the paper will be updated in this README soon. If you use this dataset, please cite… See the full description on the dataset page: https://huggingface.co/datasets/allenanie/kernelbench_with_prompts.KernelBenchX
KernelBenchX
Reproducible evaluation benchmark for Triton GPU-kernel code generation by LLMs — measures buildability, numerical correctness against a deterministic test suite, and end-to-end speedup vs. a GPU-matched golden reference.
Paper: arXiv:2605.04956 · hf.co/papers/2605.04956
Evaluation harness: https://github.com/BonnieW05/KernelBenchX
Configs
Config
Rows
What it is
tasks
176
Benchmark task specs + PyTorch reference + deterministic test harness… See the full description on the dataset page: https://huggingface.co/datasets/BonnieWang/KernelBenchX.KernelBench
KernelBench
A benchmark designed to evaluate the ability of LLMs to generate efficient GPU kernels for optimizing neural network performance
Version
[07-21-2025] This HF dataset version has been updated to v0.1
Citation
@misc{ouyang2024kernelbench,
title={KernelBench: Can LLMs Write GPU Kernels?},
author={Anne Ouyang and Simon Guo and Azalia Mirhoseini},
year={2024},
url={https://scalingintelligence.stanford.edu/blogs/kernelbench/},
}
KernelBench
Dataset Card for Dataset Name
This is the copy from Stanford's KernelBench (https://huggingface.co/datasets/ScalingIntelligence/KernelBench).
Dataset Details
Level 1: 100 Problems
Level 2: 100 Problems
Level 3: 50 Problems
Level 4: 20 Problems
Plan:
We want to try and tackle the dataset as well at MBZUAI / Imperial College London.
KernelBench
KernelBench
A benchmark designed to evaluate the ability of LLMs to generate efficient GPU kernels for optimizing neural network performance
Version
[07-21-2025] This HF dataset version has been updated to v0.1
Citation
@misc{ouyang2024kernelbench,
title={KernelBench: Can LLMs Write GPU Kernels?},
author={Anne Ouyang and Simon Guo and Azalia Mirhoseini},
year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/liushy99/KernelBench.KernelBench-bf16
KernelBench-bf16
KernelBench, with bfloat16 data type. Generated from li-plus/KernelBench.
kernel-bench
