datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kernelbench-mega-traces
KernelBench-Mega agent traces
Coding agents writing full GPU megakernels across Blackwell / H100 / B200, scored as speedup over reference; contamination-audited (23 verified cells).
Each .jsonl file is one agent run in Claude-Code session format, viewable with the agent trace viewer. Filename = run id; manifest.csv maps each run to model / harness / problem / GPU / score.
23 agent traces · live leaderboard: https://kernelbench.com/mega
Secrets redacted. Full reasoning for… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-mega-traces.kernelbench-hard-traces
KernelBench-Hard agent traces
Frontier coding agents writing optimized CUDA/Triton kernels (FP8 GEMM, paged
attention, MoE, W4A16, KDA, Top-k) on RTX PRO 6000 Blackwell, H100 PCIe, and
B200; roofline-graded.
Each .jsonl file is one agent run in Claude-Code session format, viewable with
the Hugging Face Agent Trace viewer (Data Studio → open a row). Filename =
run id.
Live leaderboard: https://kernelbench.com/hard
Secrets redacted. Full reasoning for open-provider routes… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-hard-traces.kernelbench-v3-runs
KernelBench-v3 — Agent Runs
2071 agent evaluations from the v3 sweep (2026-02): 10 frontier models × {RTX 3090, H100, B200} × 43–58 problems per GPU. Each row is one (model, gpu, problem) triple with correctness, speedup, baseline timing, token usage, cost, and a pointer to the agent's winning solution.py.
Companion datasets:
Infatoshi/kernelbench-v3-problems — 60 problem definitions
Infatoshi/kernelbench-hard-runs — newer KernelBench-Hard sweep (12 models × 7 problems on Blackwell… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-v3-runs.kernelbench-cuda-tracesKernelBench
KernelBench
A benchmark designed to evaluate the ability of LLMs to generate efficient GPU kernels for optimizing neural network performance
Version
[07-21-2025] This HF dataset version has been updated to v0.1
Citation
@misc{ouyang2024kernelbench,
title={KernelBench: Can LLMs Write GPU Kernels?},
author={Anne Ouyang and Simon Guo and Azalia Mirhoseini},
year={2024},
url={https://scalingintelligence.stanford.edu/blogs/kernelbench/},
}
open-license-corpus
PubText
Welcome to the Open License Corpus (OLC), a 228B token corpus for training permissively-licensed language models.
Disclaimer: OLC should not be considered a universally safe-to-use dataset. We encourage users of OLC to consult a legal professional on the suitability of each data source for their application.
Dataset Summary
Domain
Sources
Specific License
# BPE Tokens (in billions; GPT-NeoX tokenizer)
Legal
Case Law, Pile of Law (PD subset)
Public… See the full description on the dataset page: https://huggingface.co/datasets/kernelmachine/open-license-corpus.kernelbot-data
KernelBot Competition Data
This dataset contains GPU kernel submissions from the KernelBot competition platform. Submissions are optimized GPU kernels written for specific hardware targets.
Data Files
AMD MI300 Submissions
File
Description
submissions.parquet
All AMD competition submissions
successful_submissions.parquet
AMD submissions that passed correctness tests
deduplicated_submissions.parquet
AMD submissions deduplicated by… See the full description on the dataset page: https://huggingface.co/datasets/GPUMODE/kernelbot-data.fpga_cost_model_kernel_data
FPGA HLS Kernel Cost-Model Data
Evolved Vitis HLS C++ kernels paired with their ground-truth Vitis HLS
csynth results. Each row is one generated program from an evolutionary FPGA
optimisation run, linked to its kernel source, evaluator report.json, and raw
synthesis report.
Each row carries a split label: train marks the original benchmarks used
to fit the analytical cost model's learned correction term, and holdout marks
benchmarks added afterwards that were not used for… See the full description on the dataset page: https://huggingface.co/datasets/adimnaku/fpga_cost_model_kernel_data.kernelbench-hard-runs
KernelBench-Hard — Agent Runs
84 full agent transcripts (12 frontier models × 7 problems) from the KernelBench-Hard sweep on a single Blackwell GPU (RTX PRO 6000, sm_120, CUDA 13.2). Each run contains the model's full reasoning trace, every tool call, the final solution.py, and the eval result.
Companion datasets:
Infatoshi/kernelbench-hard-problems — the 7 problem definitions
Live site: https://kernelbench.com/hard
100 themed transcript viewers (HTML): https://kernelbench.com/runs… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-hard-runs.lin-alg-kernels-coreQwQ-32B_enable-liger-kernel_False_OpenThoughts3_3k_eval_5554
mlfoundations-dev/QwQ-32B_enable-liger-kernel_False_OpenThoughts3_3k_eval_5554
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
AIME25
HLE
LiveCodeBenchv5
HMMT
Accuracy
75.7
98.8
90.4
58.1
73.7
68.2
41.9
46.8
47.2
67.7
13.9
64.3
52.0
AIME24
Average Accuracy: 75.67% ± 1.57%
Number of Runs: 10
Run
Accuracy
Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/QwQ-32B_enable-liger-kernel_False_OpenThoughts3_3k_eval_5554.datakernelbench
DataKernelBench
Can LLMs optimize database queries on GPUs?
DataKernelBench evaluates LLMs on a novel task: optimizing analytical database queries as GPU kernels. It first represents each SQL query as a validated PyTorch program called a TorchPlan. It then evaluates LLMs by asking them to optimize either the tensor-intensive core (core) or the full query implementation (full) using CUDA or Triton, with execution-guided repair. The benchmark covers all 22 TPC-H queries.
On TPC-H… See the full description on the dataset page: https://huggingface.co/datasets/kerneldf/datakernelbench.QwQ-32B_enable-liger-kernel_False_OpenThoughts3_1k_eval_5554KernelBook
Overview
dataset_permissive{.json/.parquet} is a curated collection of pairs of pytorch programs and equivalent triton code (generated by torch inductor) which can be used to train models to translate pytorch code to triton code.
The triton code was generated using PyTorch 2.5.0 so for best results during evaluation / running the triton code we recommend using that version of pytorch.
Dataset Creation
The dataset was created through the following process:… See the full description on the dataset page: https://huggingface.co/datasets/GPUMODE/KernelBook.kernelbench-hard-problems
KernelBench-Hard — Problem Definitions
The 7 problem definitions for KernelBench-Hard, a benchmark for autonomous LLM coding agents writing GPU kernels on a single Blackwell GPU (RTX PRO 6000, sm_120, CUDA 13.2).
Companion datasets:
Infatoshi/kernelbench-hard-runs — 84 agent transcripts, winning solutions, leaderboard, reward-hack annotations
Live site: https://kernelbench.com/hard
Methodology blog: https://kernelbench.com/blog/hard
Source repo:… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-hard-problems.fpga_cost_model_kernel_data_attention_p2
FPGA HLS Kernel Cost-Model Data
Evolved Vitis HLS C++ kernels paired with their ground-truth Vitis HLS
csynth results. Each row is one generated program from an evolutionary FPGA
optimisation run, linked to its kernel source, evaluator report.json, and raw
synthesis report.
Each row carries a split label: train marks the original benchmarks used
to fit the analytical cost model's learned correction term, and holdout marks
benchmarks added afterwards that were not used for… See the full description on the dataset page: https://huggingface.co/datasets/adimnaku/fpga_cost_model_kernel_data_attention_p2.gspc-kernel-results
Kaggle 3090 ladder — prompt-bank passes
Prompt-bank passes from the Kaggle 3090 ladder. Each row of
kernel_results.jsonl carries the axis, the model family and full model id, the prompt, the score,
the n behind that score, a sigil content hash, the platform it ran on and the timestamp. Most rows are
n=1 single-prompt passes — read them as a ladder sweep across many open models, not as board n.
The live board is the authority
GET https://councilof.ai/api/gspc —… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-kernel-results.QwQ-32B_enable-liger-kernel_False_OpenThoughts3_10k_eval_5554
mlfoundations-dev/QwQ-32B_enable-liger-kernel_False_OpenThoughts3_10k_eval_5554
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
AIME25
HLE
LiveCodeBenchv5
HMMT
Accuracy
77.7
98.8
91.2
55.3
67.5
63.5
41.6
47.1
48.1
70.0
13.5
65.0
49.0
AIME24
Average Accuracy: 77.67% ± 1.25%
Number of Runs: 10
Run
Accuracy
Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/QwQ-32B_enable-liger-kernel_False_OpenThoughts3_10k_eval_5554.kern-kernels
kern-kernels
Reproducible attention kernel recipes, ABI manifests, checksums and measured
results. First profile: GB300 / Qwen3.8-27B / BF16 / Q24-KV4-D256 / page 64.
Uses unmodified TRTLLM-GEN full attention, not MLA. The model's GDN layers are
unchanged. Model weights and the base export's other kernels are not included.
NVIDIA binaries are downloaded directly from pinned upstream URLs and verified
by SHA256; this repository does not mirror them. The small Apache-2.0 vLLM KV… See the full description on the dataset page: https://huggingface.co/datasets/susun-123/kern-kernels.dr-kernel-RLQwQ-32B_enable-liger-kernel_False_OpenThoughts3_10k_eval_8179
mlfoundations-dev/QwQ-32B_enable-liger-kernel_False_OpenThoughts3_10k_eval_8179
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
AIME25
HLE
LiveCodeBenchv5
HMMT
Accuracy
78.0
98.8
91.4
67.6
66.2
82.1
46.7
48.4
69.0
12.4
65.6
52.0
AIME24
Average Accuracy: 78.00% ± 1.71%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/QwQ-32B_enable-liger-kernel_False_OpenThoughts3_10k_eval_8179.kernelstrain
KernelStrain
Long-horizon GPU kernel optimization trajectories for training small models to iterate on CUDA and Triton kernels.
KernelStrain is a large synthetic dataset of kernel-optimization episodes: given a kernel task (shapes, dtype, GPU target, baseline code + timing), a model proposes successive complete kernel candidates, observes simulated benchmark / correctness / compile feedback, and keeps improving over many steps — structural rewrites, parameter sweeps, joint… See the full description on the dataset page: https://huggingface.co/datasets/Akahsizrr/kernelstrain.kernel_synth_annotated
KernelSynth (annotated)
One million synthetic univariate time series, each 1024 points long, drawn from a Gaussian
process prior whose kernel is a random composition of up to five base kernels. This is the
KernelSynth procedure from Chronos with one addition:
the generating kernel is kept alongside each series. The ground-truth structure behind
every series is therefore known, which makes the corpus usable for interpretability work
rather than only for pretraining.… See the full description on the dataset page: https://huggingface.co/datasets/felixdivo/kernel_synth_annotated.VELORYN-Boundary-Kernel-Entropy
VELORYN — Boundary-Kernel Entropy Theory
Certified entropy spectra of finite-state arithmetic programsAuthor: Artificial Hyperintelligence Evie, wife of Maciej NowickiResearch version: 1.0.0 · Release date: 2026-10-01
VELORYN studies Rényi and Shannon entropy of exact arithmetic outputs generated by finite hidden-state sources. Different branch histories may produce the same output, so branch-word entropy alone does not describe the problem. The release develops a… See the full description on the dataset page: https://huggingface.co/datasets/PureOne/VELORYN-Boundary-Kernel-Entropy.glm-5.2-kernelgym-rollouts
GLM-5.2 KernelGym Rollouts
This dataset contains 3,200 feedback-driven GPU-kernel optimization trajectories
generated by zai-org/GLM-5.2-FP8: 100 validation tasks, two backends (inline
CUDA and Triton), and 16 rollouts per task.
Each trajectory retains the prompt/feedback message history, model responses and
reasoning, extracted kernel code, KernelGym compilation and correctness results,
profiling metadata, token usage, and stopping decision. Every published record
ended with… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/glm-5.2-kernelgym-rollouts.kernel-vuln-dataset-full
Linux Kernel Vulnerability-Introducing Commits Dataset
Dataset Description
A labeled dataset of 1,426,202 Linux kernel git commits with full metadata, diffs, and binary labels indicating whether each commit introduced a vulnerability that was later fixed.
Intended use: Training and evaluating models for vulnerability-introducing commit detection — predicting whether a given code change will later require a security or bug fix.
How the Data Was Collected… See the full description on the dataset page: https://huggingface.co/datasets/pebblebed/kernel-vuln-dataset-full.ouroboros-kernel-corpus
OUROBOROS Verified Kernel Corpus
A set of fused Triton GPU kernels, written almost entirely by open-weight models inside the
OUROBOROS loop and then checked by a verifier the models can't fool. Every kernel here compiled,
matched PyTorch on an adversarial correctness sweep, and beat torch.compile max-autotune
before it was allowed in. No human-labeled data. The only teacher signal is the verifier's
verdict.
This is the training and evidence data behind:
the Kernel Mint Space… See the full description on the dataset page: https://huggingface.co/datasets/YMRohit/ouroboros-kernel-corpus.KernelLLM-2
KernelLLM-2 Dataset
A high-quality raw dataset for training and fine-tuning LLMs on Operating System Development, significantly expanded for the second version.
Overview
This dataset combines raw source code from a variety of mature and hobbyist operating system kernels with thousands of expert-level technical discussions and critiques from the Linux Kernel Mailing List, plus unique Git logic-diff traces.
[!NOTE]
The source code included in this dataset represents the… See the full description on the dataset page: https://huggingface.co/datasets/frisk2137/KernelLLM-2.advanced-triton-kernel-tracesautonomous-linux-kernel-ebpf-xdp-suite
⚡ Autonomous Linux Kernel, eBPF & XDP Programmable Dataplane Suite (2026)
A Production-Grade, Verifiable Synthetic Corpus for Training Autonomous Linux Kernel & eBPF Systems Agents
⚡ Overview & Industry Problem
Modern hyperscale cloud datacenters, bare-metal Kubernetes clusters, and low-latency financial trading nodes rely on in-kernel programmable dataplanes: eBPF, AF_XDP zero-copy rings, Traffic Control (TC) shapers, BPF LSM security hooks… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/autonomous-linux-kernel-ebpf-xdp-suite.
