Team Ai
23 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ScalingIntelligence /kernelbench-samples KernelBench Samples Samples from experiments for KernelBench, described in our arxiv Learn more about KernelBench from our Paper Github Repo The samples are organized as such baseline_eval (Section 4 Baseline) repeated_sampling (Section 5.1.1 Repeated Sampling) iterative_refinement (Section 5.1.2 Iterative Refinement of Generations) Within each folder, we organize the results by /level/model/problem_{id}/sample_{id}. The inner most .json file contains the generated kernel and… See the full description on the dataset page: https://huggingface.co/datasets/ScalingIntelligence/kernelbench-samples.3 likes11k downloads2y agoHugging Face02Infatoshi /kernelbench-mega-traces KernelBench-Mega agent traces Coding agents writing full GPU megakernels across Blackwell / H100 / B200, scored as speedup over reference; contamination-audited (23 verified cells). Each .jsonl file is one agent run in Claude-Code session format, viewable with the agent trace viewer. Filename = run id; manifest.csv maps each run to model / harness / problem / GPU / score. 23 agent traces · live leaderboard: https://kernelbench.com/mega Secrets redacted. Full reasoning for… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-mega-traces.tabularn<1K18 likes6.5k downloads7h agoHugging Face03Infatoshi /kernelbench-hard-traces KernelBench-Hard agent traces Frontier coding agents writing optimized CUDA/Triton kernels (FP8 GEMM, paged attention, MoE, W4A16, KDA, Top-k) on RTX PRO 6000 Blackwell, H100 PCIe, and B200; roofline-graded. Each .jsonl file is one agent run in Claude-Code session format, viewable with the Hugging Face Agent Trace viewer (Data Studio → open a row). Filename = run id. Live leaderboard: https://kernelbench.com/hard Secrets redacted. Full reasoning for open-provider routes… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-hard-traces.tabulartext-generationn<1K16 likes6k downloads13d agoHugging Face04Infatoshi /kernelbench-v3-runs KernelBench-v3 — Agent Runs 2071 agent evaluations from the v3 sweep (2026-02): 10 frontier models × {RTX 3090, H100, B200} × 43–58 problems per GPU. Each row is one (model, gpu, problem) triple with correctness, speedup, baseline timing, token usage, cost, and a pointer to the agent's winning solution.py. Companion datasets: Infatoshi/kernelbench-v3-problems — 60 problem definitions Infatoshi/kernelbench-hard-runs — newer KernelBench-Hard sweep (12 models × 7 problems on Blackwell… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-v3-runs.tabular1K<n<10K3 likes2.4k downloads5mo agoHugging Face05Infatoshi /kernelbench-cuda-tracestabularn<1K2 likes1.9k downloads7h agoHugging Face06ScalingIntelligence /KernelBench KernelBench A benchmark designed to evaluate the ability of LLMs to generate efficient GPU kernels for optimizing neural network performance Version [07-21-2025] This HF dataset version has been updated to v0.1 Citation @misc{ouyang2024kernelbench, title={KernelBench: Can LLMs Write GPU Kernels?}, author={Anne Ouyang and Simon Guo and Azalia Mirhoseini}, year={2024}, url={https://scalingintelligence.stanford.edu/blogs/kernelbench/}, } tabularn<1K51 likes1.8k downloads1y agoHugging Face07Infatoshi /kernelbench-hard-runs KernelBench-Hard — Agent Runs 84 full agent transcripts (12 frontier models × 7 problems) from the KernelBench-Hard sweep on a single Blackwell GPU (RTX PRO 6000, sm_120, CUDA 13.2). Each run contains the model's full reasoning trace, every tool call, the final solution.py, and the eval result. Companion datasets: Infatoshi/kernelbench-hard-problems — the 7 problem definitions Live site: https://kernelbench.com/hard 100 themed transcript viewers (HTML): https://kernelbench.com/runs… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-hard-runs.tabularn<1K3 likes589 downloads5mo agoHugging Face08Infatoshi /kernelbench-v3-problems KernelBench-v3 — Problem Definitions The full set of problem definitions for KernelBench-v3 — the previous-generation sweep (2026-02) covering 10 frontier models across 3 NVIDIA GPUs (RTX 3090, H100, B200), with 43–58 problems per GPU. Companion datasets: Infatoshi/kernelbench-v3-runs — 2071 eval rows + winning agent solutions Infatoshi/kernelbench-hard-problems — the newer KernelBench-Hard suite (single-Blackwell, 7 problems, 12 models) Live site: https://kernelbench.com/v3 Source… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-v3-problems.n<1K0 likes383 downloads5mo agoHugging Face09lukeleeai /kernelbench-rag-content0 likes344 downloads8mo agoHugging Face10Elfsong /KernelBench-M KernelBench-M The measurement artifact for Measuring the Checker: Mutation Analysis for GPU-Kernel Benchmark Oracles: the mutation operators, the verified CUDA substrates they mutate, the kill witnesses, and the pipeline that produced every number in the paper. Layout rules/ 124 mutation rules, six families (mutator.py loads all of them) substrates/ 208 gate-verified CUDA implementations, one per KernelBench problem: the mutation… See the full description on the dataset page: https://huggingface.co/datasets/Elfsong/KernelBench-M.10K<n<100K1 likes334 downloads1mo agoHugging Face11Infatoshi /kernelbench-hard-problems KernelBench-Hard — Problem Definitions The 7 problem definitions for KernelBench-Hard, a benchmark for autonomous LLM coding agents writing GPU kernels on a single Blackwell GPU (RTX PRO 6000, sm_120, CUDA 13.2). Companion datasets: Infatoshi/kernelbench-hard-runs — 84 agent transcripts, winning solutions, leaderboard, reward-hack annotations Live site: https://kernelbench.com/hard Methodology blog: https://kernelbench.com/blog/hard Source repo:… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-hard-problems.textn<1K0 likes200 downloads5mo agoHugging Face12allenanie /kernelbench_with_promptsThis is a version of KernelBench where the prompts to produce the Triton and cuda kernel are explicitly saved in the JSON data files. It only contains Level 1, 2, 3 kernels. The prompt is the same as what is provided in the original KernelBench repo. The dataset is prepared by Jiin Woo during her internship at AWS Annapurna Labs, the lab behind Trainium chips. This dataset is part of an unreleased paper, and the paper will be updated in this README soon. If you use this dataset, please cite… See the full description on the dataset page: https://huggingface.co/datasets/allenanie/kernelbench_with_prompts.tabularn<1K1 likes81 downloads1y agoHugging Face13BonnieWang /KernelBenchX KernelBenchX Reproducible evaluation benchmark for Triton GPU-kernel code generation by LLMs — measures buildability, numerical correctness against a deterministic test suite, and end-to-end speedup vs. a GPU-matched golden reference. Paper: arXiv:2605.04956 · hf.co/papers/2605.04956 Evaluation harness: https://github.com/BonnieW05/KernelBenchX Configs Config Rows What it is tasks 176 Benchmark task specs + PyTorch reference + deterministic test harness… See the full description on the dataset page: https://huggingface.co/datasets/BonnieWang/KernelBenchX.tabulartext-generationn<1K2 likes60 downloads5mo agoHugging Face14andrew-wang /kernelbench_harness_expansiontext1K<n<10K0 likes29 downloads10mo agoHugging Face15Infatoshi /kernelbench-hard-submissions KernelBench-Hard - Agent Kernel Submissions Real CUDA / Triton GPU kernels written autonomously by frontier coding models on KernelBench-Hard: each model gets one unlimited-time autonomous run per problem to write the fastest kernel it can for an NVIDIA RTX PRO 6000 Blackwell (SM120), graded as peak_fraction of the hardware roofline. This is the unlimited-time generation (June 2026): 8 frontier models (Claude Opus 4.8, GPT-5.5, GLM-5.2, MiniMax-M3, Gemini 3.5 Flash, Kimi… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-hard-submissions.texttext-generationn<1K1 likes29 downloads4mo agoHugging Face16mzweilin /KernelBench KernelBench A benchmark designed to evaluate the ability of LLMs to generate efficient GPU kernels for optimizing neural network performance Version [07-21-2025] This HF dataset version has been updated to v0.1 Citation @misc{ouyang2024kernelbench, title={KernelBench: Can LLMs Write GPU Kernels?}, author={Anne Ouyang and Simon Guo and Azalia Mirhoseini}, year={2024}, url={https://scalingintelligence.stanford.edu/blogs/kernelbench/}, } tabularn<1K0 likes25 downloads8mo agoHugging Face17ai-nikolai /KernelBench Dataset Card for Dataset Name This is the copy from Stanford's KernelBench (https://huggingface.co/datasets/ScalingIntelligence/KernelBench). Dataset Details Level 1: 100 Problems Level 2: 100 Problems Level 3: 50 Problems Level 4: 20 Problems Plan: We want to try and tackle the dataset as well at MBZUAI / Imperial College London. tabularn<1K1 likes22 downloads1y agoHugging Face18liushy99 /KernelBench KernelBench A benchmark designed to evaluate the ability of LLMs to generate efficient GPU kernels for optimizing neural network performance Version [07-21-2025] This HF dataset version has been updated to v0.1 Citation @misc{ouyang2024kernelbench, title={KernelBench: Can LLMs Write GPU Kernels?}, author={Anne Ouyang and Simon Guo and Azalia Mirhoseini}, year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/liushy99/KernelBench.tabularn<1K0 likes21 downloads3mo agoHugging Face19li-plus /KernelBench-bf16 KernelBench-bf16 KernelBench, with bfloat16 data type. Generated from li-plus/KernelBench. tabularn<1K0 likes20 downloads1y agoHugging Face20justus27 /kernelbench-genesystextn<1K0 likes18 downloads1y agoHugging Face21jiliu1 /kernelbench_hiptext1K<n<10K0 likes14 downloads5mo agoHugging Face22yueqis /kernelbench_doc0 likes8 downloads1y agoHugging Face23dtadpole /kernel-benchtabularn<1K0 likes7 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.