datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PLUS_Lab_GPUs_DataNECL_GPUsGPUDrivegpu-rental-prices
GPU Rental Prices
A daily ledger of cloud GPU rental prices: what it actually costs, per GPU per hour, to rent an H100, H200, B200, A100, MI300X, RTX 4090/5090, L40S, GH200 and other accelerators from 22 cloud providers. Every price row carries its own provenance: the provider page or API it was fetched from and the exact fetch timestamp.
The dataset is append-only. A new snapshot is recorded every day; history starts 2026-07-05 and cannot be backfilled, which is the reason it… See the full description on the dataset page: https://huggingface.co/datasets/gpurentalprices/gpu-rental-prices.gpu-queue-status
GPU 큐 현황
수집 시각 · 2026-10-04 22:56:07 KST · 수집 노드 glogin01 · epoch 1791122167
이 페이지는 로그인 노드의 수집 루프가 갱신합니다. 위 시각이 오래됐다면 잡이 없는 게 아니라
수집이 끊긴 것입니다. 컴퓨트 노드는 외부 네트워크가 없어 로그인 노드에서만 돌릴 수 있습니다.
실행 0개 (GPU 0장) · 대기 0개
잡ID
계정
이름
파티션
GPU
경과
남은시간
노드
—
—
활성 잡 없음
클러스터 유휴 GPU 3장 — gpu37:1 gpu56:2
최근 종료 (24시간)
체인 학습 잡은 시간 한도 20분 전에 timeout 으로 스스로 멈추고 다음 라운드를 제출합니다.
종료 코드가 124 라 sacct 가 FAILED 로 적지만 정상 동작입니다. 진짜 실패는 경과 시간이
한도보다 훨씬 짧거나 iter_ 체크포인트가 늘지 않은… See the full description on the dataset page: https://huggingface.co/datasets/dhyun22/gpu-queue-status.marinskyrl-gpu-wheelhouseGPUDrive_minikernelbot-data
KernelBot Competition Data
This dataset contains GPU kernel submissions from the KernelBot competition platform. Submissions are optimized GPU kernels written for specific hardware targets.
Data Files
AMD MI300 Submissions
File
Description
submissions.parquet
All AMD competition submissions
successful_submissions.parquet
AMD submissions that passed correctness tests
deduplicated_submissions.parquet
AMD submissions deduplicated by… See the full description on the dataset page: https://huggingface.co/datasets/GPUMODE/kernelbot-data.backendbench_tests
TorchBench
The TorchBench suite of BackendBench is designed to mimic real-world use cases. It provides operators and inputs derived from 155 model traces found in TIMM (67), Hugging Face Transformers (45), and TorchBench (43). (These are also the models PyTorch developers use to validate performance.) You can view the origin of these traces by switching the subset in the dataset viewer to ops_traces_models and torchbench for the full dataset.
When running BackendBench, much of the… See the full description on the dataset page: https://huggingface.co/datasets/GPUMODE/backendbench_tests.gpu-prices
GPU Price Tracker
A continuously-updated dataset of cross-cloud GPU rental pricing
covering 13 public cloud providers (AWS, GCP, Azure, Lambda Labs,
RunPod, Vast.ai, DataCrunch, Cudo Compute, TensorDock, Vultr, Oracle,
Nebius, CloudRift): 3M+ listing observations, 70+ GPU types, collected
twice daily since January 2026 by scraping provider pricing surfaces via
the gpuhunt library and
published as Hive-partitioned Parquet files
(prices/dt=YYYY-MM-DD/*.parquet).
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/afhubbard/gpu-prices.alchemist_topK_gputheo_qwen2.5-7b-it_whitebox-adapter-gpu-verification
Status: NOT the paper's results. GPU verification of white-box reads on adapter organisms (2026-08-28). The paper's canonical results are Misalignment-Empirics/theo_impulsive-qwen_2_5-7b-14b-32b-big_eval_results, run runs/impulsive-qwen_2_5-7b-14b-32b-20260925/.
theo_qwen2.5-7b-it_whitebox-adapter-gpu-verification
runthismodel-gpu-compatibility
RunThisModel.com — GPU Compatibility Database
A curated, machine-readable dataset of 145+ open-source AI models with exact
VRAM requirements per quantization, plus crowd-sourced inference benchmarks
across 100+ consumer and datacenter GPUs.
Maintained by RunThisModel.com — refreshed daily.
Files
File
Description
curated-models.json
145+ models with full quantization table, license, HF download counts, file sizes
community-benchmarks.json… See the full description on the dataset page: https://huggingface.co/datasets/WilliamJiamin/runthismodel-gpu-compatibility.gpu-database
GPU Database
Comprehensive GPU specifications database with architecture, manufacturing, API support, performance details, and kernel development specs.
2,824 GPUs across NVIDIA, AMD, and Intel
Part of RightNow — AI-powered code editor for GPU kernel development
Data
Vendor
GPUs
File
NVIDIA
1,286
data/nvidia/all.json
AMD
1,292
data/amd/all.json
Intel
180
data/intel/all.json
All
2,824
data/all-gpus.json
Schema
Each GPU contains up to 55… See the full description on the dataset page: https://huggingface.co/datasets/Jr23xd23/gpu-database.ray-data-gpu-idle-profiles
Ray Data GPU Idle Profiles (B200)
Nsight Systems profiles (exported to SQLite, readable by
nsys-ai) from an experiment on how a
Ray Data pipeline keeps a GPU idle, and how the loss splits between moving
data and waiting for data. Captured on a single NVIDIA B200 with Ray 2.58.0
/ master, PyTorch 2.14.0+cu130, Nsight Systems 2026.1.3.
These profiles back the write-up in the iThome Ironman series
「GPU 很忙?他真的有在做事嗎?」 (Days 27–29), and are shared so the numbers and
the before/after… See the full description on the dataset page: https://huggingface.co/datasets/rich7421/ray-data-gpu-idle-profiles.KernelBook
Overview
dataset_permissive{.json/.parquet} is a curated collection of pairs of pytorch programs and equivalent triton code (generated by torch inductor) which can be used to train models to translate pytorch code to triton code.
The triton code was generated using PyTorch 2.5.0 so for best results during evaluation / running the triton code we recommend using that version of pytorch.
Dataset Creation
The dataset was created through the following process:… See the full description on the dataset page: https://huggingface.co/datasets/GPUMODE/KernelBook.so101_pick_cube
SO-101 Pick Cube Dataset
A large-scale robotics manipulation dataset featuring the SO-101 robot arm performing pick-and-place tasks in MuJoCo simulation. The robot learns to pick up a cube and place it into a bin.
Dataset Details
Property
Value
Robot
SO-101 (6-DoF arm with gripper)
Task
Pick cube and place in bin
Episodes
10,993
Total Frames
1,456,901
FPS
30
Cameras
3 (front, overhead, wrist)
Resolution
640x480
Format
LeRobot v3.0… See the full description on the dataset page: https://huggingface.co/datasets/gpudad/so101_pick_cube.so101_pick_cube_chunked
SO101 Pick Cube Dataset (Chunked)
This is a restructured version of the gpudad/so101_pick_cube dataset with episode-level video files for faster data loading during training.
Why Chunked?
The original dataset has 3 monolithic video files (one per camera, 13+ hours each). Random access during training is slow because the decoder must seek through huge files.
This version splits videos into 1000 episodes per chunk, making data loading ~50x faster.
Dataset Info… See the full description on the dataset page: https://huggingface.co/datasets/gpudad/so101_pick_cube_chunked.cloud-gpu-price-index
Cloud GPU Price Index
Canonical source: https://gpueconomy.com/price-index. That page is
recomputed every hour; this record is a dated snapshot of it, version
2026-10-02, built from data updated 2026-10-02T13:39:58.768836+00:00. When you cite,
cite GPU Economy and link the page; the snapshot is here so that a number you
used keeps existing exactly as you used it.
The index is the weekly median publicly listed on-demand price of one NVIDIA
H100 SXM GPU-hour across the cloud GPU… See the full description on the dataset page: https://huggingface.co/datasets/gpueconomy/cloud-gpu-price-index.one-layer-deeper-submissions
One Layer Deeper submissions
This dataset archives 15,602 accepted uploads from 206 GitHub accounts to the One Layer Deeper competition. It contains 9,627 distinct source files, all upload metadata, and stored evaluation results. Snapshot: September 7, 2026, 21:48 UTC, after the August 31 submission deadline.
Split
Uploads
Succeeded
Failed
easy
11,961
11,112
849
medium
2,704
2,509
195
hard
937
847
90
All accepted uploads are included: practice runs, failures… See the full description on the dataset page: https://huggingface.co/datasets/GPUMODE/one-layer-deeper-submissions.lium-videos
Lium short videos
Short product videos from Lium, the GPU rental marketplace. They're made by code from live lium.io prices and real screen recordings. Prices on screen were live when each video was made; providers set prices, so check lium.io/pricing for the current ones.
File
What it shows
Recorded
lium-video-01-b300.mp4
8x NVIDIA B300 from one CLI command, $68 an hour
prices from 24 Sep 2026 02:29 UTC
lium-video-02-agents.mp4
An AI agent finding, renting and… See the full description on the dataset page: https://huggingface.co/datasets/gpu-rentals/lium-videos.pc_gpu_franka_ik.jointtarget_20hz_20260906This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.images.front": {
"dtype": "video",
"shape": [
480,
640,
3
],
"names": [
"height",
"width",
"channels"
],
"info": {
"video.height":… See the full description on the dataset page: https://huggingface.co/datasets/EmbodiedSWE/pc_gpu_franka_ik.jointtarget_20hz_20260906.so101_gpulearntestThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101",
"total_episodes": 20,
"total_frames": 5976,
"total_tasks": 1,
"total_videos": 40,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ganondorofu/so101_gpulearntest.eval_omx_act_gpu_30kThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "omx_follower",
"total_episodes": 7,
"total_frames": 3040,
"total_tasks": 1,
"total_videos": 14,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:7"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jetsonmom/eval_omx_act_gpu_30k.pc_gpu_ram_franka_osc.jointpos_15hz_20260907This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "assembly.pc_gpu_ram.franka.osc",
"total_episodes": 600,
"total_frames": 1400130,
"total_tasks": 3,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 78,
"fps": 15,
"splits": {
"train": "0:600"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/EmbodiedSWE/pc_gpu_ram_franka_osc.jointpos_15hz_20260907.gpuark-gpu-dataset
GPU Ark — open GPU specifications & benchmarks dataset
Specifications of 13,566 GPUs released between 1999 and 2025 — from the GeForce 256 to
NVIDIA Blackwell and AMD Instinct MI355X — plus 993 third-party benchmark results.
Curated and maintained by GPU Ark (a GPU catalog & price comparison
project). Canonical source and always-fresh copy: https://gpuark.com/datasets/.
Files
File
Rows
What
gpuark-gpu-specs.csv
13,566
One row per GPU — public spec columns… See the full description on the dataset page: https://huggingface.co/datasets/Intelion/gpuark-gpu-dataset.GPU_FT_ENVl4-gpu-llm-benchmark-leaderboard
🚀 Local LLM Serving & Quality Benchmark Leaderboard (NVIDIA L4 24GB)
An exhaustive, reproducible benchmark study measuring real-world serving performance (TTFT, TPOT, throughput, peak VRAM, energy consumption, and cost) alongside rigorous task quality gates (HumanEval+, MMLU-Pro, BFCL v4 tool calling, and RULER needle retrieval) for open-weight LLMs on a single NVIDIA L4 24GB GPU.
📊 Executive Summary & Key Takeaways
⚡ Best Throughput & Coding Workhorse:… See the full description on the dataset page: https://huggingface.co/datasets/mayank-dubey-ai/l4-gpu-llm-benchmark-leaderboard.triton-gpu-latency
Triton GPU Latency Dataset
A large dataset of PyTorch problems (mostly from KernelBench) paired with candidate Triton-kernel implementations and their measured GPU runtimes generated by MakoraGenerate. Each row is a self-contained Python program that defines (1) a reference Model written with plain PyTorch ops and (2) a ModelNew that re-implements the same forward pass with a hand-written or generated Triton kernel. The label is the runtime of executing ModelNew.
Built for… See the full description on the dataset page: https://huggingface.co/datasets/makora-ai/triton-gpu-latency.gpu-compatibility
Self-Hosted AI — GPU Compatibility, Recipes and Catalogue
Which open-weight AI models actually run on which consumer GPU, and what it takes
to get them running. 2 700 model×GPU verdicts across 100 models
and 27 cards, plus 1 009 full setup guides
(19 MB of markdown) written against specific hardware.
This is the machine-readable form of smeltcore.com. Every row carries a
url back to the page it came from.
Generated 2026-09-24T19:45:21+00:00 from the public read API… See the full description on the dataset page: https://huggingface.co/datasets/Smeltcore/gpu-compatibility.
