Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ScalingIntelligence /kernelbench-samples KernelBench Samples Samples from experiments for KernelBench, described in our arxiv Learn more about KernelBench from our Paper Github Repo The samples are organized as such baseline_eval (Section 4 Baseline) repeated_sampling (Section 5.1.1 Repeated Sampling) iterative_refinement (Section 5.1.2 Iterative Refinement of Generations) Within each folder, we organize the results by /level/model/problem_{id}/sample_{id}. The inner most .json file contains the generated kernel and… See the full description on the dataset page: https://huggingface.co/datasets/ScalingIntelligence/kernelbench-samples.3 likes11k downloads2y agoHugging Face02Infatoshi /kernelbench-mega-traces KernelBench-Mega agent traces Coding agents writing full GPU megakernels across Blackwell / H100 / B200, scored as speedup over reference; contamination-audited (23 verified cells). Each .jsonl file is one agent run in Claude-Code session format, viewable with the agent trace viewer. Filename = run id; manifest.csv maps each run to model / harness / problem / GPU / score. 23 agent traces · live leaderboard: https://kernelbench.com/mega Secrets redacted. Full reasoning for… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-mega-traces.tabularn<1K18 likes6.5k downloads15h agoHugging Face03Infatoshi /kernelbench-hard-traces KernelBench-Hard agent traces Frontier coding agents writing optimized CUDA/Triton kernels (FP8 GEMM, paged attention, MoE, W4A16, KDA, Top-k) on RTX PRO 6000 Blackwell, H100 PCIe, and B200; roofline-graded. Each .jsonl file is one agent run in Claude-Code session format, viewable with the Hugging Face Agent Trace viewer (Data Studio → open a row). Filename = run id. Live leaderboard: https://kernelbench.com/hard Secrets redacted. Full reasoning for open-provider routes… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-hard-traces.tabulartext-generationn<1K16 likes6k downloads13d agoHugging Face04kernelmachine /open-license-corpus PubText Welcome to the Open License Corpus (OLC), a 228B token corpus for training permissively-licensed language models. Disclaimer: OLC should not be considered a universally safe-to-use dataset. We encourage users of OLC to consult a legal professional on the suitability of each data source for their application. Dataset Summary Domain Sources Specific License # BPE Tokens (in billions; GPT-NeoX tokenizer) Legal Case Law, Pile of Law (PD subset) Public… See the full description on the dataset page: https://huggingface.co/datasets/kernelmachine/open-license-corpus.texttext-generation10M<n<100M19 likes4.5k downloads3y agoHugging Face05Infatoshi /kernelbench-v3-runs KernelBench-v3 — Agent Runs 2071 agent evaluations from the v3 sweep (2026-02): 10 frontier models × {RTX 3090, H100, B200} × 43–58 problems per GPU. Each row is one (model, gpu, problem) triple with correctness, speedup, baseline timing, token usage, cost, and a pointer to the agent's winning solution.py. Companion datasets: Infatoshi/kernelbench-v3-problems — 60 problem definitions Infatoshi/kernelbench-hard-runs — newer KernelBench-Hard sweep (12 models × 7 problems on Blackwell… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-v3-runs.tabular1K<n<10K3 likes2.4k downloads5mo agoHugging Face06Infatoshi /kernelbench-cuda-tracestabularn<1K2 likes1.9k downloads15h agoHugging Face07ScalingIntelligence /KernelBench KernelBench A benchmark designed to evaluate the ability of LLMs to generate efficient GPU kernels for optimizing neural network performance Version [07-21-2025] This HF dataset version has been updated to v0.1 Citation @misc{ouyang2024kernelbench, title={KernelBench: Can LLMs Write GPU Kernels?}, author={Anne Ouyang and Simon Guo and Azalia Mirhoseini}, year={2024}, url={https://scalingintelligence.stanford.edu/blogs/kernelbench/}, } tabularn<1K51 likes1.8k downloads1y agoHugging Face08GPUMODE /kernelbot-data KernelBot Competition Data This dataset contains GPU kernel submissions from the KernelBot competition platform. Submissions are optimized GPU kernels written for specific hardware targets. Data Files AMD MI300 Submissions File Description submissions.parquet All AMD competition submissions successful_submissions.parquet AMD submissions that passed correctness tests deduplicated_submissions.parquet AMD submissions deduplicated by… See the full description on the dataset page: https://huggingface.co/datasets/GPUMODE/kernelbot-data.tabular100K<n<1M50 likes1.2k downloads2mo agoHugging Face09P2SAMAPA /p2-etf-kernel-conditional-moment-results2 likes814 downloads17d agoHugging Face10adimnaku /fpga_cost_model_kernel_data FPGA HLS Kernel Cost-Model Data Evolved Vitis HLS C++ kernels paired with their ground-truth Vitis HLS csynth results. Each row is one generated program from an evolutionary FPGA optimisation run, linked to its kernel source, evaluator report.json, and raw synthesis report. Each row carries a split label: train marks the original benchmarks used to fit the analytical cost model's learned correction term, and holdout marks benchmarks added afterwards that were not used for… See the full description on the dataset page: https://huggingface.co/datasets/adimnaku/fpga_cost_model_kernel_data.tabulartabular-regressionn<1K0 likes627 downloads3mo agoHugging Face11muahmed7338 /kernelascent-tasks KernelAscent — public dev split KernelAscent is a benchmark for recursive self-improvement (RSI): a model optimizes the GPU kernels used to train itself, and we measure whether kernel-optimization capability compounds across rounds. This is the public dev split, released for self-benchmarking and research; the leaderboard is scored on a private held-out split. Project & code: https://github.com/ahmd-mohsin/KernelAscent Leaderboard & docs:… See the full description on the dataset page: https://huggingface.co/datasets/muahmed7338/kernelascent-tasks.text-generationn<1K0 likes624 downloads19d agoHugging Face12TokenBender /lin-alg-kernels-coretextn<1K0 likes617 downloads4mo agoHugging Face13Infatoshi /kernelbench-hard-runs KernelBench-Hard — Agent Runs 84 full agent transcripts (12 frontier models × 7 problems) from the KernelBench-Hard sweep on a single Blackwell GPU (RTX PRO 6000, sm_120, CUDA 13.2). Each run contains the model's full reasoning trace, every tool call, the final solution.py, and the eval result. Companion datasets: Infatoshi/kernelbench-hard-problems — the 7 problem definitions Live site: https://kernelbench.com/hard 100 themed transcript viewers (HTML): https://kernelbench.com/runs… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-hard-runs.tabularn<1K3 likes589 downloads5mo agoHugging Face14Infatoshi /kernelbench-v3-problems KernelBench-v3 — Problem Definitions The full set of problem definitions for KernelBench-v3 — the previous-generation sweep (2026-02) covering 10 frontier models across 3 NVIDIA GPUs (RTX 3090, H100, B200), with 43–58 problems per GPU. Companion datasets: Infatoshi/kernelbench-v3-runs — 2071 eval rows + winning agent solutions Infatoshi/kernelbench-hard-problems — the newer KernelBench-Hard suite (single-Blackwell, 7 problems, 12 models) Live site: https://kernelbench.com/v3 Source… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-v3-problems.n<1K0 likes383 downloads5mo agoHugging Face15lukeleeai /kernelbench-rag-content0 likes344 downloads8mo agoHugging Face16Elfsong /KernelBench-M KernelBench-M The measurement artifact for Measuring the Checker: Mutation Analysis for GPU-Kernel Benchmark Oracles: the mutation operators, the verified CUDA substrates they mutate, the kill witnesses, and the pipeline that produced every number in the paper. Layout rules/ 124 mutation rules, six families (mutator.py loads all of them) substrates/ 208 gate-verified CUDA implementations, one per KernelBench problem: the mutation… See the full description on the dataset page: https://huggingface.co/datasets/Elfsong/KernelBench-M.10K<n<100K1 likes334 downloads1mo agoHugging Face17mlfoundations-dev /QwQ-32B_enable-liger-kernel_False_OpenThoughts3_3k_eval_5554 mlfoundations-dev/QwQ-32B_enable-liger-kernel_False_OpenThoughts3_3k_eval_5554 Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces AIME25 HLE LiveCodeBenchv5 HMMT Accuracy 75.7 98.8 90.4 58.1 73.7 68.2 41.9 46.8 47.2 67.7 13.9 64.3 52.0 AIME24 Average Accuracy: 75.67% ± 1.57% Number of Runs: 10 Run Accuracy Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/QwQ-32B_enable-liger-kernel_False_OpenThoughts3_3k_eval_5554.tabular10K<n<100K0 likes322 downloads1y agoHugging Face18mlfoundations-dev /QwQ-32B_enable-liger-kernel_False_OpenThoughts3_1k_eval_5554tabular10K<n<100K0 likes313 downloads1y agoHugging Face19adimnaku /fpga_cost_model_kernel_data_attention_p2 FPGA HLS Kernel Cost-Model Data Evolved Vitis HLS C++ kernels paired with their ground-truth Vitis HLS csynth results. Each row is one generated program from an evolutionary FPGA optimisation run, linked to its kernel source, evaluator report.json, and raw synthesis report. Each row carries a split label: train marks the original benchmarks used to fit the analytical cost model's learned correction term, and holdout marks benchmarks added afterwards that were not used for… See the full description on the dataset page: https://huggingface.co/datasets/adimnaku/fpga_cost_model_kernel_data_attention_p2.tabulartabular-regressionn<1K0 likes274 downloads3mo agoHugging Face20kerneldf /datakernelbench DataKernelBench Can LLMs optimize database queries on GPUs? DataKernelBench evaluates LLMs on a novel task: optimizing analytical database queries as GPU kernels. It first represents each SQL query as a validated PyTorch program called a TorchPlan. It then evaluates LLMs by asking them to optimize either the tensor-intensive core (core) or the full query implementation (full) using CUDA or Triton, with execution-guided repair. The benchmark covers all 22 TPC-H queries. On TPC-H… See the full description on the dataset page: https://huggingface.co/datasets/kerneldf/datakernelbench.texttext-generationn<1K1 likes268 downloads1mo agoHugging Face21muahmed7338 /kernelascent KernelAscent — public dev split KernelAscent is a benchmark for recursive self-improvement (RSI): a model optimizes the GPU kernels used to train itself, and we measure whether kernel-optimization capability compounds across rounds. This is the public dev split, released for self-benchmarking and research; the leaderboard is scored on a private held-out split. Project & code: https://github.com/ahmd-mohsin/KernelAscent Leaderboard & docs:… See the full description on the dataset page: https://huggingface.co/datasets/muahmed7338/kernelascent.text-generationn<1K0 likes257 downloads19d agoHugging Face22GPUMODE /KernelBook Overview dataset_permissive{.json/.parquet} is a curated collection of pairs of pytorch programs and equivalent triton code (generated by torch inductor) which can be used to train models to translate pytorch code to triton code. The triton code was generated using PyTorch 2.5.0 so for best results during evaluation / running the triton code we recommend using that version of pytorch. Dataset Creation The dataset was created through the following process:… See the full description on the dataset page: https://huggingface.co/datasets/GPUMODE/KernelBook.tabular10K<n<100K58 likes256 downloads4mo agoHugging Face23rtferraz /cuda-kernel-engineering CUDA Kernel Engineering — Portfolio A hands-on CUDA kernel engineering portfolio built on an NVIDIA L4 GPU (GCP). Covers the complete path from first kernel to research-backed hypotheses that were empirically falsified, with Nsight Compute profiling evidence at every step. Each project teaches a specific optimization, measures its impact against cuBLAS, and documents both positive and negative results. Hardware: NVIDIA L4 (sm_89, 300 GB/s, 23 GB GDDR6)Stack: CUDA 12.4 (nvcc) /… See the full description on the dataset page: https://huggingface.co/datasets/rtferraz/cuda-kernel-engineering.0 likes231 downloads5mo agoHugging Face24ewedubs /linux-kernel-commits-aireason-instruct Linux Kernel Code Patches Dataset High-quality Linux kernel commit patches for training code generation and understanding models. Dataset Description This dataset contains 144,089 curated Linux kernel commits with: Commit messages (instruction) Smart-extracted code context (input) Unified diff patches (output) Optional AI quality scores and reasoning Dataset Variants Variant Examples Description super_ultra 206 AI-recommended commits (Reasoning… See the full description on the dataset page: https://huggingface.co/datasets/ewedubs/linux-kernel-commits-aireason-instruct.10K<n<100K1 likes222 downloads10mo agoHugging Face25ADAPT-Chase /titans-memory-kernels Titans Memory Kernel - Nova Lineage Evolved memory kernels for the Titans Memory system. Structure nova_prime.safetensors - The Prime kernel (16M params, 4096 dim) nova_v1.safetensors - Production-hardened V1 (identity preserved) REGISTRY.jsonl - Lineage registry with metadata variants/ - 137 generation snapshots from evolution Parameters Dimensions: 4096 Memory Size: ~65MB per kernel Evolution: Hardened with decay=0.99999, lr=0.0001 Usage from… See the full description on the dataset page: https://huggingface.co/datasets/ADAPT-Chase/titans-memory-kernels.0 likes211 downloads9mo agoHugging Face26csoai /gspc-kernel-results Kaggle 3090 ladder — prompt-bank passes Prompt-bank passes from the Kaggle 3090 ladder. Each row of kernel_results.jsonl carries the axis, the model family and full model id, the prompt, the score, the n behind that score, a sigil content hash, the platform it ran on and the timestamp. Most rows are n=1 single-prompt passes — read them as a ladder sweep across many open models, not as board n. The live board is the authority GET https://councilof.ai/api/gspc —… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-kernel-results.tabularother1K<n<10K0 likes201 downloads2d agoHugging Face27Infatoshi /kernelbench-hard-problems KernelBench-Hard — Problem Definitions The 7 problem definitions for KernelBench-Hard, a benchmark for autonomous LLM coding agents writing GPU kernels on a single Blackwell GPU (RTX PRO 6000, sm_120, CUDA 13.2). Companion datasets: Infatoshi/kernelbench-hard-runs — 84 agent transcripts, winning solutions, leaderboard, reward-hack annotations Live site: https://kernelbench.com/hard Methodology blog: https://kernelbench.com/blog/hard Source repo:… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-hard-problems.textn<1K0 likes200 downloads5mo agoHugging Face28kernel-14 /SemanticAlign-Bench SemanticAlign-Bench A benchmark for evaluating AI agents on structured claim extraction from top-tier ML conference papers. Each paper is decomposed into Semantic Alignment Units (SAU) — atomic, self-contained implementation propositions — across four diagnostic dimensions spanning numerical precision to pipeline-level workflow. Agents are evaluated on whether they can reproduce these claims without hallucination, omission, or misordering. The Four SAU Dimensions… See the full description on the dataset page: https://huggingface.co/datasets/kernel-14/SemanticAlign-Bench.imagequestion-answering1K<n<10K1 likes182 downloads5mo agoHugging Face29susun-123 /kern-kernels kern-kernels Reproducible attention kernel recipes, ABI manifests, checksums and measured results. First profile: GB300 / Qwen3.8-27B / BF16 / Q24-KV4-D256 / page 64. Uses unmodified TRTLLM-GEN full attention, not MLA. The model's GDN layers are unchanged. Model weights and the base export's other kernels are not included. NVIDIA binaries are downloaded directly from pinned upstream URLs and verified by SHA256; this repository does not mirror them. The small Apache-2.0 vLLM KV… See the full description on the dataset page: https://huggingface.co/datasets/susun-123/kern-kernels.tabularn<1K0 likes164 downloads1mo agoHugging Face30AhNr /dr-kernel-RLtexttext-generation1K<n<10K3 likes157 downloads27d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.