Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01uclanecl /NECL_GPUstabularn<1K10 likes8.9k downloads4h agoHugging Face02GPUMODE /kernelbot-data KernelBot Competition Data This dataset contains GPU kernel submissions from the KernelBot competition platform. Submissions are optimized GPU kernels written for specific hardware targets. Data Files AMD MI300 Submissions File Description submissions.parquet All AMD competition submissions successful_submissions.parquet AMD submissions that passed correctness tests deduplicated_submissions.parquet AMD submissions deduplicated by… See the full description on the dataset page: https://huggingface.co/datasets/GPUMODE/kernelbot-data.tabular100K<n<1M50 likes1.1k downloads2mo agoHugging Face03afhubbard /gpu-prices GPU Price Tracker A continuously-updated dataset of cross-cloud GPU rental pricing covering 13 public cloud providers (AWS, GCP, Azure, Lambda Labs, RunPod, Vast.ai, DataCrunch, Cudo Compute, TensorDock, Vultr, Oracle, Nebius, CloudRift): 3M+ listing observations, 70+ GPU types, collected twice daily since January 2026 by scraping provider pricing surfaces via the gpuhunt library and published as Hive-partitioned Parquet files (prices/dt=YYYY-MM-DD/*.parquet). The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/afhubbard/gpu-prices.tabulartabular-regression1M<n<10M0 likes850 downloads12h agoHugging Face04GPUMODE /backendbench_tests TorchBench The TorchBench suite of BackendBench is designed to mimic real-world use cases. It provides operators and inputs derived from 155 model traces found in TIMM (67), Hugging Face Transformers (45), and TorchBench (43). (These are also the models PyTorch developers use to validate performance.) You can view the origin of these traces by switching the subset in the dataset viewer to ops_traces_models and torchbench for the full dataset. When running BackendBench, much of the… See the full description on the dataset page: https://huggingface.co/datasets/GPUMODE/backendbench_tests.tabular10K<n<100K4 likes805 downloads1y agoHugging Face05Jr23xd23 /gpu-database GPU Database Comprehensive GPU specifications database with architecture, manufacturing, API support, performance details, and kernel development specs. 2,824 GPUs across NVIDIA, AMD, and Intel Part of RightNow — AI-powered code editor for GPU kernel development Data Vendor GPUs File NVIDIA 1,286 data/nvidia/all.json AMD 1,292 data/amd/all.json Intel 180 data/intel/all.json All 2,824 data/all-gpus.json Schema Each GPU contains up to 55… See the full description on the dataset page: https://huggingface.co/datasets/Jr23xd23/gpu-database.tabulartext-generation1K<n<10K1 likes260 downloads9mo agoHugging Face06GPUMODE /KernelBook Overview dataset_permissive{.json/.parquet} is a curated collection of pairs of pytorch programs and equivalent triton code (generated by torch inductor) which can be used to train models to translate pytorch code to triton code. The triton code was generated using PyTorch 2.5.0 so for best results during evaluation / running the triton code we recommend using that version of pytorch. Dataset Creation The dataset was created through the following process:… See the full description on the dataset page: https://huggingface.co/datasets/GPUMODE/KernelBook.tabular10K<n<100K58 likes255 downloads4mo agoHugging Face07gpueconomy /cloud-gpu-price-index Cloud GPU Price Index Canonical source: https://gpueconomy.com/price-index. That page is recomputed every hour; this record is a dated snapshot of it, version 2026-10-02, built from data updated 2026-10-02T13:39:58.768836+00:00. When you cite, cite GPU Economy and link the page; the snapshot is here so that a number you used keeps existing exactly as you used it. The index is the weekly median publicly listed on-demand price of one NVIDIA H100 SXM GPU-hour across the cloud GPU… See the full description on the dataset page: https://huggingface.co/datasets/gpueconomy/cloud-gpu-price-index.tabularn<1K0 likes190 downloads8d agoHugging Face08GPUMODE /one-layer-deeper-submissions One Layer Deeper submissions This dataset archives 15,602 accepted uploads from 206 GitHub accounts to the One Layer Deeper competition. It contains 9,627 distinct source files, all upload metadata, and stored evaluation results. Snapshot: September 7, 2026, 21:48 UTC, after the August 31 submission deadline. Split Uploads Succeeded Failed easy 11,961 11,112 849 medium 2,704 2,509 195 hard 937 847 90 All accepted uploads are included: practice runs, failures… See the full description on the dataset page: https://huggingface.co/datasets/GPUMODE/one-layer-deeper-submissions.tabular10K<n<100K1 likes173 downloads1mo agoHugging Face09mayank-dubey-ai /l4-gpu-llm-benchmark-leaderboard 🚀 Local LLM Serving & Quality Benchmark Leaderboard (NVIDIA L4 24GB) An exhaustive, reproducible benchmark study measuring real-world serving performance (TTFT, TPOT, throughput, peak VRAM, energy consumption, and cost) alongside rigorous task quality gates (HumanEval+, MMLU-Pro, BFCL v4 tool calling, and RULER needle retrieval) for open-weight LLMs on a single NVIDIA L4 24GB GPU. 📊 Executive Summary & Key Takeaways ⚡ Best Throughput & Coding Workhorse:… See the full description on the dataset page: https://huggingface.co/datasets/mayank-dubey-ai/l4-gpu-llm-benchmark-leaderboard.tabulartext-generationn<1K0 likes138 downloads2mo agoHugging Face10Intelion /gpuark-gpu-dataset GPU Ark — open GPU specifications & benchmarks dataset Specifications of 13,566 GPUs released between 1999 and 2025 — from the GeForce 256 to NVIDIA Blackwell and AMD Instinct MI355X — plus 993 third-party benchmark results. Curated and maintained by GPU Ark (a GPU catalog & price comparison project). Canonical source and always-fresh copy: https://gpuark.com/datasets/. Files File Rows What gpuark-gpu-specs.csv 13,566 One row per GPU — public spec columns… See the full description on the dataset page: https://huggingface.co/datasets/Intelion/gpuark-gpu-dataset.tabular10K<n<100K0 likes127 downloads5mo agoHugging Face11Smeltcore /gpu-compatibility Self-Hosted AI — GPU Compatibility, Recipes and Catalogue Which open-weight AI models actually run on which consumer GPU, and what it takes to get them running. 2 700 model×GPU verdicts across 100 models and 27 cards, plus 1 009 full setup guides (19 MB of markdown) written against specific hardware. This is the machine-readable form of smeltcore.com. Every row carries a url back to the page it came from. Generated 2026-09-24T19:45:21+00:00 from the public read API… See the full description on the dataset page: https://huggingface.co/datasets/Smeltcore/gpu-compatibility.tabulartable-question-answering1K<n<10K0 likes123 downloads16d agoHugging Face12makora-ai /triton-gpu-latency Triton GPU Latency Dataset A large dataset of PyTorch problems (mostly from KernelBench) paired with candidate Triton-kernel implementations and their measured GPU runtimes generated by MakoraGenerate. Each row is a self-contained Python program that defines (1) a reference Model written with plain PyTorch ops and (2) a ModelNew that re-implements the same forward pass with a hand-written or generated Triton kernel. The label is the runtime of executing ModelNew. Built for… See the full description on the dataset page: https://huggingface.co/datasets/makora-ai/triton-gpu-latency.texttext-generation100K<n<1M17 likes119 downloads4mo agoHugging Face13codezakh /gpu-forecasters-puct-search-eventsCompanion artifact for GPU Forecasters: Language Models as Selective Surrogates for Kernel Runtime Optimization. Code: codezakh/gpu-surrogates. Events emitted during the kernel searches in this work. You can reconstruct everything that happened in a search from these events: which kernels were tried, in what order, with what runtimes. See the kernel-search code at codezakh/gpu-surrogates for the payload schema of each event kind. Loading from datasets import load_dataset #… See the full description on the dataset page: https://huggingface.co/datasets/codezakh/gpu-forecasters-puct-search-events.tabular10K<n<100K0 likes110 downloads4mo agoHugging Face14codezakh /gpu-forecasters-eval-setCompanion artifact for GPU Forecasters: Language Models as Selective Surrogates for Kernel Runtime Optimization. Code: codezakh/gpu-surrogates. Held-out evaluation set used in the paper. Each row is one (reference, candidate) kernel pair on a GPU Mode task, with the candidate's measured speedup over the reference on an A100. Loading from datasets import load_dataset # all six packs combined ds = load_dataset("codezakh/gpu-forecasters-eval-set", name="combined", split="eval") #… See the full description on the dataset page: https://huggingface.co/datasets/codezakh/gpu-forecasters-eval-set.tabularn<1K0 likes106 downloads4mo agoHugging Face15referencesource /gpu-cuda-pytorch-compatibility PyTorch release, CUDA build and NVIDIA driver pairings Canonical, always-current version: https://referencesource.org/gpu-cuda-pytorch-compatibility/ Machine-readable: https://referencesource.org/gpu-cuda-pytorch-compatibility/data.json — this mirror is a point-in-time copy. Last verified: 2026-09-29 Stale after: 2026-12-28 (past this date, prefer the canonical copy — it re-verifies on a cadence this snapshot does not) Records: 45 One record per (PyTorch release, CUDA build)… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/gpu-cuda-pytorch-compatibility.textn<1K0 likes98 downloads4d agoHugging Face16beatsprom /autonomous-cloud-gpu-slurm-serving-suite ⚡ Autonomous Cloud GPU Infrastructure, Slurm Orchestration & Distributed Serving Suite (2026) A Production-Grade, Verifiable Synthetic Corpus for Training Autonomous AI Supercomputing & LLM Serving Agents ⚡ Overview & Industry Problem Operating massive AI supercomputers (thousands of NVIDIA H100/H200 and Blackwell GPUs) requires coordinating Slurm cluster schedules, topology-aware NVLink cliques, NCCL AllReduce rings, RoCE v2 lossless fabrics… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/autonomous-cloud-gpu-slurm-serving-suite.tabulartext-generation1K<n<10K0 likes96 downloads23d agoHugging Face17beatsprom /autonomous-gpu-kernel-triton-cuda-suite-2026 ⚡ Autonomous GPU Kernel, Triton & CUDA Architecture Suite (2026) A Production-Grade, Verifiable Synthetic Corpus for Training Frontier Coding Models (Qwen 3.8, DeepSeek-V3, Llama 3.3) ⚡ Overview & Industry Problem Modern deep learning accelerators, custom ASICs, and high-performance computing clusters demand specialized, autonomous GPU kernel infrastructure: OpenAI Triton fused kernels, FlashAttention-3 forward/backward online softmax… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/autonomous-gpu-kernel-triton-cuda-suite-2026.tabulartext-generation10K<n<100K0 likes87 downloads19d agoHugging Face18beatsprom /cuda-triton-gpu-kernels-2026 ⚡ Complete 2026 CUDA & OpenAI Triton High-Performance GPU Kernel Engineering SFT/DPO Suite The definitive, production-grade synthetic alignment dataset engineered for training and fine-tuning open-weights Large Language Models (Qwen 2.5 Coder, DeepSeek-Coder, Llama 3.1) on ultra-high-throughput GPU kernel programming: NVIDIA Hopper H100 / Blackwell B200 TMA async transfers, OpenAI Triton 3.1+ FlashAttention-3, 32-bank conflict elimination, and low-bit FP8 / INT4 GEMM… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/cuda-triton-gpu-kernels-2026.tabulartext-generation1K<n<10K0 likes81 downloads1mo agoHugging Face19winglian /gpu-mode-triton-augment-sample Dataset Card for gpu-mode-triton-augment-sample This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/winglian/gpu-mode-triton-augment-sample/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/winglian/gpu-mode-triton-augment-sample.textn<1K0 likes79 downloads2y agoHugging Face20fastgpu /cloud-gpu-prices Cloud GPU Rental Prices (FastGPU) What it costs to rent a GPU in the cloud, per GPU per hour, across marketplaces, neoclouds and hyperscalers. Mirrored here once a day from fastgpu.co/dataset, where the same files are served live and every price links to the provider it came from. Browse the prices themselves at fastgpu.co. A free, openly licensed dataset of GPU cloud rental prices across the whole market: marketplaces, neoclouds, and hyperscalers. It contains a normalized… See the full description on the dataset page: https://huggingface.co/datasets/fastgpu/cloud-gpu-prices.tabulartime-series-forecasting10K<n<100K0 likes71 downloads5h agoHugging Face21referencesource /nvidia-cuda-compute-capability-by-gpu NVIDIA GPU CUDA compute capability by model Canonical, always-current version: https://referencesource.org/nvidia-cuda-compute-capability-by-gpu/ Machine-readable: https://referencesource.org/nvidia-cuda-compute-capability-by-gpu/data.json — this mirror is a point-in-time copy. Last verified: 2026-10-01 Stale after: 2027-01-29 (past this date, prefer the canonical copy — it re-verifies on a cadence this snapshot does not) Records: 442 Which CUDA compute capability version… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/nvidia-cuda-compute-capability-by-gpu.textn<1K0 likes69 downloads5d agoHugging Face22EvanOLeary /pallasbench-robust-gpu-a100 PallasBench: Robust Pallas GPU Kernel Benchmark (A100) 39/45 kernels passing on NVIDIA A100 80GB -- the first GPU-focused evaluation of JAX Pallas kernels. What is this? PallasBench is a suite of 45 JAX Pallas kernels across 3 difficulty levels. The original kernels were designed for TPU and failed on GPU because Pallas compiles to Triton on NVIDIA hardware, which has strict block size limits that TPU's Mosaic compiler does not. We fixed all 45 kernels for GPU… See the full description on the dataset page: https://huggingface.co/datasets/EvanOLeary/pallasbench-robust-gpu-a100.tabulartext-generationn<1K1 likes64 downloads4mo agoHugging Face23Swapnil007-Curious /kanvas-gpu-failure KANVAS GPU Task Failure Features Derived data for the KANVAS project: an interpretable Kolmogorov-Arnold Additive Model (KAAM) that predicts GPU cluster task failure. Code, notebooks and tests: https://github.com/Swapnil007-Curious/KANVAS Files File Rows Contents kanvas_feature_matrix.csv 18,890 10 scheduling and telemetry features plus the failure label, one row per task kanvas_phi_curves.csv 2,000 Learned per-feature curves phi_i(x) of the canonical… See the full description on the dataset page: https://huggingface.co/datasets/Swapnil007-Curious/kanvas-gpu-failure.tabulartabular-classification10K<n<100K1 likes53 downloads11d agoHugging Face24codezakh /gpu-forecasters-eval-set-predictionsCompanion artifact for GPU Forecasters: Language Models as Selective Surrogates for Kernel Runtime Optimization. Code: codezakh/gpu-surrogates. Surrogate predictions on the held-out evaluation set. Each row is one forecast from one (surrogate, repeat) on one row of codezakh/gpu-forecasters-eval-set. Loading from datasets import load_dataset # all surrogates and repeats ds = load_dataset("codezakh/gpu-forecasters-eval-set-predictions", name="combined", split="predictions") # one… See the full description on the dataset page: https://huggingface.co/datasets/codezakh/gpu-forecasters-eval-set-predictions.tabular10K<n<100K0 likes48 downloads4mo agoHugging Face25ehyo /GPU-Resources-Estimation-for-Deep-Learning-Training-Tasks GPUMemNet and GPUUtilNet Dataset This dataset accompanies the paper “GPU Memory and Utilization Estimation for Training-Aware Resource Management: Opportunities and Limitations.” It contains synthetic deep learning training configurations and their measured GPU memory consumption and utilization characteristics. Dataset configurations The dataset is divided into separate configurations because MLP, CNN, and Transformer workloads use different feature schemas.… See the full description on the dataset page: https://huggingface.co/datasets/ehyo/GPU-Resources-Estimation-for-Deep-Learning-Training-Tasks.tabulartabular-regression10K<n<100K0 likes47 downloads4mo agoHugging Face26thsysmfh /gpu-spot-rental-prices GPU spot-rental prices Weekly medians of vast.ai marketplace asks, normalised to USD per GPU-hour (dph_total / num_gpus), for H100, H200, B200, B300, A100 — on-demand and interruptible tiers where listed. Method: cheapest-500 asks per SKU, so the series is deliberately low-biased; min / p25 / median per tier. Machine-generated from vast.ai's public marketplace API. No internal figures, no vendor quotes. data/prices.csv — the series, one row per GPU x tier per collection date… See the full description on the dataset page: https://huggingface.co/datasets/thsysmfh/gpu-spot-rental-prices.tabularn<1K0 likes47 downloads2mo agoHugging Face27AdityaMayukhSom /MixSub-LLaMA-3.2-Entities-Overlap-GPU-Scoretabular1K<n<10K0 likes46 downloads2y agoHugging Face28dipankarsarkar /gpuemu-corpus gpuemu Kernel-Correctness Corpus + Reproducibility Artifact The 26-op corpus, experiment drivers, and analysis scripts behind the four gpuemu preprints, led by "The Correctness Illusion in LLM-Generated GPU Kernels" (arXiv:2606.20128). A controlled set for measuring whether a correctness oracle actually catches the bugs LLM-generated GPU kernels routinely contain — plus the full harness that produces every table and figure in the papers. Papers it backs P1 — The… See the full description on the dataset page: https://huggingface.co/datasets/dipankarsarkar/gpuemu-corpus.texttext-classificationn<1K2 likes46 downloads4mo agoHugging Face29ohdoking /energy_consumption_by_model_and_gputabular10K<n<100K0 likes45 downloads1y agoHugging Face30GPUburnout /hcc-array-supplementtext10K<n<100K0 likes44 downloads25d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.