Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Infatoshi /kernelbench-mega-traces KernelBench-Mega agent traces Coding agents writing full GPU megakernels across Blackwell / H100 / B200, scored as speedup over reference; contamination-audited (23 verified cells). Each .jsonl file is one agent run in Claude-Code session format, viewable with the agent trace viewer. Filename = run id; manifest.csv maps each run to model / harness / problem / GPU / score. 23 agent traces · live leaderboard: https://kernelbench.com/mega Secrets redacted. Full reasoning for… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-mega-traces.tabularn<1K19 likes6.5k downloads1d agoHugging Face02Infatoshi /kernelbench-hard-traces KernelBench-Hard agent traces Frontier coding agents writing optimized CUDA/Triton kernels (FP8 GEMM, paged attention, MoE, W4A16, KDA, Top-k) on RTX PRO 6000 Blackwell, H100 PCIe, and B200; roofline-graded. Each .jsonl file is one agent run in Claude-Code session format, viewable with the Hugging Face Agent Trace viewer (Data Studio → open a row). Filename = run id. Live leaderboard: https://kernelbench.com/hard Secrets redacted. Full reasoning for open-provider routes… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-hard-traces.tabulartext-generationn<1K16 likes6k downloads13d agoHugging Face03Infatoshi /kernelbench-v3-runs KernelBench-v3 — Agent Runs 2071 agent evaluations from the v3 sweep (2026-02): 10 frontier models × {RTX 3090, H100, B200} × 43–58 problems per GPU. Each row is one (model, gpu, problem) triple with correctness, speedup, baseline timing, token usage, cost, and a pointer to the agent's winning solution.py. Companion datasets: Infatoshi/kernelbench-v3-problems — 60 problem definitions Infatoshi/kernelbench-hard-runs — newer KernelBench-Hard sweep (12 models × 7 problems on Blackwell… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-v3-runs.tabular1K<n<10K3 likes2.4k downloads5mo agoHugging Face04Infatoshi /kernelbench-cuda-tracestabularn<1K2 likes1.9k downloads1d agoHugging Face05ScalingIntelligence /KernelBench KernelBench A benchmark designed to evaluate the ability of LLMs to generate efficient GPU kernels for optimizing neural network performance Version [07-21-2025] This HF dataset version has been updated to v0.1 Citation @misc{ouyang2024kernelbench, title={KernelBench: Can LLMs Write GPU Kernels?}, author={Anne Ouyang and Simon Guo and Azalia Mirhoseini}, year={2024}, url={https://scalingintelligence.stanford.edu/blogs/kernelbench/}, } tabularn<1K51 likes1.8k downloads1y agoHugging Face06GPUMODE /kernelbot-data KernelBot Competition Data This dataset contains GPU kernel submissions from the KernelBot competition platform. Submissions are optimized GPU kernels written for specific hardware targets. Data Files AMD MI300 Submissions File Description submissions.parquet All AMD competition submissions successful_submissions.parquet AMD submissions that passed correctness tests deduplicated_submissions.parquet AMD submissions deduplicated by… See the full description on the dataset page: https://huggingface.co/datasets/GPUMODE/kernelbot-data.tabular100K<n<1M50 likes1.2k downloads2mo agoHugging Face07adimnaku /fpga_cost_model_kernel_data FPGA HLS Kernel Cost-Model Data Evolved Vitis HLS C++ kernels paired with their ground-truth Vitis HLS csynth results. Each row is one generated program from an evolutionary FPGA optimisation run, linked to its kernel source, evaluator report.json, and raw synthesis report. Each row carries a split label: train marks the original benchmarks used to fit the analytical cost model's learned correction term, and holdout marks benchmarks added afterwards that were not used for… See the full description on the dataset page: https://huggingface.co/datasets/adimnaku/fpga_cost_model_kernel_data.tabulartabular-regressionn<1K0 likes627 downloads3mo agoHugging Face08Infatoshi /kernelbench-hard-runs KernelBench-Hard — Agent Runs 84 full agent transcripts (12 frontier models × 7 problems) from the KernelBench-Hard sweep on a single Blackwell GPU (RTX PRO 6000, sm_120, CUDA 13.2). Each run contains the model's full reasoning trace, every tool call, the final solution.py, and the eval result. Companion datasets: Infatoshi/kernelbench-hard-problems — the 7 problem definitions Live site: https://kernelbench.com/hard 100 themed transcript viewers (HTML): https://kernelbench.com/runs… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-hard-runs.tabularn<1K3 likes589 downloads5mo agoHugging Face09mlfoundations-dev /QwQ-32B_enable-liger-kernel_False_OpenThoughts3_3k_eval_5554 mlfoundations-dev/QwQ-32B_enable-liger-kernel_False_OpenThoughts3_3k_eval_5554 Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces AIME25 HLE LiveCodeBenchv5 HMMT Accuracy 75.7 98.8 90.4 58.1 73.7 68.2 41.9 46.8 47.2 67.7 13.9 64.3 52.0 AIME24 Average Accuracy: 75.67% ± 1.57% Number of Runs: 10 Run Accuracy Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/QwQ-32B_enable-liger-kernel_False_OpenThoughts3_3k_eval_5554.tabular10K<n<100K0 likes322 downloads1y agoHugging Face10mlfoundations-dev /QwQ-32B_enable-liger-kernel_False_OpenThoughts3_1k_eval_5554tabular10K<n<100K0 likes313 downloads1y agoHugging Face11adimnaku /fpga_cost_model_kernel_data_attention_p2 FPGA HLS Kernel Cost-Model Data Evolved Vitis HLS C++ kernels paired with their ground-truth Vitis HLS csynth results. Each row is one generated program from an evolutionary FPGA optimisation run, linked to its kernel source, evaluator report.json, and raw synthesis report. Each row carries a split label: train marks the original benchmarks used to fit the analytical cost model's learned correction term, and holdout marks benchmarks added afterwards that were not used for… See the full description on the dataset page: https://huggingface.co/datasets/adimnaku/fpga_cost_model_kernel_data_attention_p2.tabulartabular-regressionn<1K0 likes274 downloads3mo agoHugging Face12GPUMODE /KernelBook Overview dataset_permissive{.json/.parquet} is a curated collection of pairs of pytorch programs and equivalent triton code (generated by torch inductor) which can be used to train models to translate pytorch code to triton code. The triton code was generated using PyTorch 2.5.0 so for best results during evaluation / running the triton code we recommend using that version of pytorch. Dataset Creation The dataset was created through the following process:… See the full description on the dataset page: https://huggingface.co/datasets/GPUMODE/KernelBook.tabular10K<n<100K58 likes256 downloads4mo agoHugging Face13csoai /gspc-kernel-results Kaggle 3090 ladder — prompt-bank passes Prompt-bank passes from the Kaggle 3090 ladder. Each row of kernel_results.jsonl carries the axis, the model family and full model id, the prompt, the score, the n behind that score, a sigil content hash, the platform it ran on and the timestamp. Most rows are n=1 single-prompt passes — read them as a ladder sweep across many open models, not as board n. The live board is the authority GET https://councilof.ai/api/gspc —… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-kernel-results.tabularother1K<n<10K0 likes201 downloads2d agoHugging Face14susun-123 /kern-kernels kern-kernels Reproducible attention kernel recipes, ABI manifests, checksums and measured results. First profile: GB300 / Qwen3.8-27B / BF16 / Q24-KV4-D256 / page 64. Uses unmodified TRTLLM-GEN full attention, not MLA. The model's GDN layers are unchanged. Model weights and the base export's other kernels are not included. NVIDIA binaries are downloaded directly from pinned upstream URLs and verified by SHA256; this repository does not mirror them. The small Apache-2.0 vLLM KV… See the full description on the dataset page: https://huggingface.co/datasets/susun-123/kern-kernels.tabularn<1K0 likes164 downloads1mo agoHugging Face15mlfoundations-dev /QwQ-32B_enable-liger-kernel_False_OpenThoughts3_10k_eval_5554 mlfoundations-dev/QwQ-32B_enable-liger-kernel_False_OpenThoughts3_10k_eval_5554 Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces AIME25 HLE LiveCodeBenchv5 HMMT Accuracy 77.7 98.8 91.2 55.3 67.5 63.5 41.6 47.1 48.1 70.0 13.5 65.0 49.0 AIME24 Average Accuracy: 77.67% ± 1.25% Number of Runs: 10 Run Accuracy Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/QwQ-32B_enable-liger-kernel_False_OpenThoughts3_10k_eval_5554.tabular10K<n<100K0 likes156 downloads1y agoHugging Face16mlfoundations-dev /QwQ-32B_enable-liger-kernel_False_OpenThoughts3_10k_eval_8179 mlfoundations-dev/QwQ-32B_enable-liger-kernel_False_OpenThoughts3_10k_eval_8179 Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces AIME25 HLE LiveCodeBenchv5 HMMT Accuracy 78.0 98.8 91.4 67.6 66.2 82.1 46.7 48.4 69.0 12.4 65.6 52.0 AIME24 Average Accuracy: 78.00% ± 1.71% Number of Runs: 10 Run Accuracy Questions Solved Total… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/QwQ-32B_enable-liger-kernel_False_OpenThoughts3_10k_eval_8179.tabular10K<n<100K0 likes143 downloads1y agoHugging Face17beatsprom /cuda-triton-gpu-kernels-2026 ⚡ Complete 2026 CUDA & OpenAI Triton High-Performance GPU Kernel Engineering SFT/DPO Suite The definitive, production-grade synthetic alignment dataset engineered for training and fine-tuning open-weights Large Language Models (Qwen 2.5 Coder, DeepSeek-Coder, Llama 3.1) on ultra-high-throughput GPU kernel programming: NVIDIA Hopper H100 / Blackwell B200 TMA async transfers, OpenAI Triton 3.1+ FlashAttention-3, 32-bank conflict elimination, and low-bit FP8 / INT4 GEMM… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/cuda-triton-gpu-kernels-2026.tabulartext-generation1K<n<10K0 likes100 downloads28d agoHugging Face18marin-community /glm-5.2-kernelgym-rollouts GLM-5.2 KernelGym Rollouts This dataset contains 3,200 feedback-driven GPU-kernel optimization trajectories generated by zai-org/GLM-5.2-FP8: 100 validation tasks, two backends (inline CUDA and Triton), and 16 rollouts per task. Each trajectory retains the prompt/feedback message history, model responses and reasoning, extracted kernel code, KernelGym compilation and correctness results, profiling metadata, token usage, and stopping decision. Every published record ended with… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/glm-5.2-kernelgym-rollouts.tabulartext-generation1K<n<10K2 likes98 downloads2mo agoHugging Face19mlfoundations-dev /QwQ-32B_enable-liger-kernel_False_OpenThoughts3_3k_eval_8179 mlfoundations-dev/QwQ-32B_enable-liger-kernel_False_OpenThoughts3_3k_eval_8179 Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces AIME25 HLE LiveCodeBenchv5 HMMT Accuracy 79.3 99.2 90.8 73.1 65.7 83.8 47.5 48.4 66.0 13.0 65.9 51.7 AIME24 Average Accuracy: 79.33% ± 1.23% Number of Runs: 10 Run Accuracy Questions Solved Total… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/QwQ-32B_enable-liger-kernel_False_OpenThoughts3_3k_eval_8179.tabular10K<n<100K0 likes92 downloads1y agoHugging Face20beatsprom /autonomous-linux-kernel-ebpf-xdp-suite ⚡ Autonomous Linux Kernel, eBPF & XDP Programmable Dataplane Suite (2026) A Production-Grade, Verifiable Synthetic Corpus for Training Autonomous Linux Kernel & eBPF Systems Agents ⚡ Overview & Industry Problem Modern hyperscale cloud datacenters, bare-metal Kubernetes clusters, and low-latency financial trading nodes rely on in-kernel programmable dataplanes: eBPF, AF_XDP zero-copy rings, Traffic Control (TC) shapers, BPF LSM security hooks… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/autonomous-linux-kernel-ebpf-xdp-suite.tabulartext-generation1K<n<10K0 likes92 downloads20d agoHugging Face21quguanni /kernel-vuln-dataset-full Linux Kernel Vulnerability-Introducing Commits Dataset Dataset Description A labeled dataset of 1,426,202 Linux kernel git commits with full metadata, diffs, and binary labels indicating whether each commit introduced a vulnerability that was later fixed. Intended use: Training and evaluating models for vulnerability-introducing commit detection — predicting whether a given code change will later require a security or bug fix. How the Data Was Collected… See the full description on the dataset page: https://huggingface.co/datasets/quguanni/kernel-vuln-dataset-full.tabulartext-classification1M<n<10M1 likes87 downloads8mo agoHugging Face22adimnaku /fpga_cost_model_kernel_data_mamba_p2 FPGA HLS Kernel Cost-Model Data Evolved Vitis HLS C++ kernels paired with their ground-truth Vitis HLS csynth results. Each row is one generated program from an evolutionary FPGA optimisation run, linked to its kernel source, evaluator report.json, and raw synthesis report. Each row carries a split label: train marks the original benchmarks used to fit the analytical cost model's learned correction term, and holdout marks benchmarks added afterwards that were not used for… See the full description on the dataset page: https://huggingface.co/datasets/adimnaku/fpga_cost_model_kernel_data_mamba_p2.tabulartabular-regressionn<1K0 likes86 downloads3mo agoHugging Face23allenanie /kernelbench_with_promptsThis is a version of KernelBench where the prompts to produce the Triton and cuda kernel are explicitly saved in the JSON data files. It only contains Level 1, 2, 3 kernels. The prompt is the same as what is provided in the original KernelBench repo. The dataset is prepared by Jiin Woo during her internship at AWS Annapurna Labs, the lab behind Trainium chips. This dataset is part of an unreleased paper, and the paper will be updated in this README soon. If you use this dataset, please cite… See the full description on the dataset page: https://huggingface.co/datasets/allenanie/kernelbench_with_prompts.tabularn<1K1 likes81 downloads1y agoHugging Face24pebblebed /kernel-vuln-dataset-full Linux Kernel Vulnerability-Introducing Commits Dataset Dataset Description A labeled dataset of 1,426,202 Linux kernel git commits with full metadata, diffs, and binary labels indicating whether each commit introduced a vulnerability that was later fixed. Intended use: Training and evaluating models for vulnerability-introducing commit detection — predicting whether a given code change will later require a security or bug fix. How the Data Was Collected… See the full description on the dataset page: https://huggingface.co/datasets/pebblebed/kernel-vuln-dataset-full.tabulartext-classification1M<n<10M2 likes80 downloads8mo agoHugging Face25pjt222 /ga104-cuda-kernels GA104 Hand-Optimized CUDA Kernel Corpus A measurement corpus of hand-optimized CUDA / SASS kernels targeting the RTX 3070 Ti (GA104, sm_86, Ampere). Every kernel is written without cuBLAS, cuDNN, or PyTorch in the optimized path; vendor libraries are linked only for measured comparison under kernels/reference/. This dataset is for SASS and GPU-optimization researchers — it pairs each .cu source with its compiled machine code and its disassembly, so the exact instruction stream a… See the full description on the dataset page: https://huggingface.co/datasets/pjt222/ga104-cuda-kernels.tabularn<1K0 likes80 downloads5mo agoHugging Face26adimnaku /fpga_cost_model_kernel_data_qwen3_mlp FPGA HLS Kernel Cost-Model Data Evolved Vitis HLS C++ kernels paired with their ground-truth Vitis HLS csynth results. Each row is one generated program from an evolutionary FPGA optimisation run, linked to its kernel source, evaluator report.json, and raw synthesis report. Each row carries a split label: train marks the original benchmarks used to fit the analytical cost model's learned correction term, and holdout marks benchmarks added afterwards that were not used for… See the full description on the dataset page: https://huggingface.co/datasets/adimnaku/fpga_cost_model_kernel_data_qwen3_mlp.tabulartabular-regressionn<1K0 likes80 downloads3mo agoHugging Face27beatsprom /autonomous-gpu-kernel-triton-cuda-suite-2026 ⚡ Autonomous GPU Kernel, Triton & CUDA Architecture Suite (2026) A Production-Grade, Verifiable Synthetic Corpus for Training Frontier Coding Models (Qwen 3.8, DeepSeek-V3, Llama 3.3) ⚡ Overview & Industry Problem Modern deep learning accelerators, custom ASICs, and high-performance computing clusters demand specialized, autonomous GPU kernel infrastructure: OpenAI Triton fused kernels, FlashAttention-3 forward/backward online softmax… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/autonomous-gpu-kernel-triton-cuda-suite-2026.tabulartext-generation10K<n<100K0 likes72 downloads17d agoHugging Face28siro1 /kernelbook-glm4-evalstabular10K<n<100K0 likes70 downloads9mo agoHugging Face29GAIR /daVinci-kernel-sfttabular10K<n<100K0 likes70 downloads4mo agoHugging Face30willychan21 /ParallelKernelBench_Kernels ParallelKernelBench Kernels Net-new multi-GPU CUDA kernels generated by LLMs for ParallelKernelBench. Each subdirectory under solutions/ is one model run. File names match the benchmark problem stems (e.g. 17_rope_allgather_cuda.py ↔ problem 17_rope_allgather in willychan21/ParallelKernelBench_Problems). Layout solutions/ <run_id>/ <stem>_cuda.py ... Runs (1 run(s), 87 kernel files) run_id kernels path… See the full description on the dataset page: https://huggingface.co/datasets/willychan21/ParallelKernelBench_Kernels.tabulartext-generationn<1K0 likes64 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.