datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LESA-FLUX-A100-cutpoint-RTX4090-20260928
A100-origin cutpoint follow-up
This is a new research campaign from the pinned public A100 artifact revision 45671fe9ea2bd97491d921338cbf85a853e34578, not a continuation of the lost fresh-RTX4090 2k checkpoints. The origin teacher cache and base model are referenced by revision/SHA rather than copied here. New measurements and any new checkpoints/derived tensors will be uploaded by stage. Until verified result folders appear, this repository contains only the plan, not completed… See the full description on the dataset page: https://huggingface.co/datasets/LamTNguyen/LESA-FLUX-A100-cutpoint-RTX4090-20260928.GenEmotions-RTX5070Tirtx-5090-benchmarks
RTX 5090 LLM Benchmarks
Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig.
Quality Benchmarks
Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency.
Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/rtx-5090-benchmarks.multigpu-beam-rtx3060-scaling-20261004
Completed four-series scaling measurement
Full report · Machine-readable results · CSV table · Article fragment · Experiment protocol
We measured MultiGPU Beam Search on a single host with eight identical RTX 3060 12 GiB GPUs, using 1, 2, 4 and 8 devices. The fixed-work experiment used a global beam of 4,194,304 and twelve complete search steps, expanding 637,616,736 actions in every run. With the common execution profile, median synchronized end-to-end search times were 85.107… See the full description on the dataset page: https://huggingface.co/datasets/TryDotAtwo/multigpu-beam-rtx3060-scaling-20261004.qwen3-coder-gb10-vs-rtx5090-benchmark
NVIDIA GB10 vs. GeForce RTX 5090 - Local LLM Inference Benchmark
Model: Qwen3-Coder-30B-A3B-InstructFormat: GGUF, Q4_K_M, 18.63 GBRuntime: LM Studio / llama.cppAuthor: Efehan A.Benchmark date: 5 August 2026
This repository contains a decode-focused local inference benchmark comparing an NVIDIA GB10 system with a Windows workstation containing two GeForce RTX 5090 GPUs. Telemetry shows that the inference workload was carried primarily by a single RTX 5090 (GPU 0), while GPU 1… See the full description on the dataset page: https://huggingface.co/datasets/mreltera/qwen3-coder-gb10-vs-rtx5090-benchmark.rtx5090-energy-benchmark
RTX 5090 LLM Energy Benchmark
First energy efficiency benchmark of 4-bit quantization on NVIDIA RTX 5090 (Blackwell architecture).
Key Finding
4-bit quantization increases energy consumption by up to 29% for models < 5B parameters.
The crossover point where quantization becomes beneficial is ~5B parameters.
Results
Model
FP16 Energy
4-bit Energy
Change
TinyLlama 1.1B
1,659 J/1k
2,098 J/1k
+26.5% 🔴
Qwen2 1.5B
2,411 J/1k
3,120 J/1k
+29.4% 🔴… See the full description on the dataset page: https://huggingface.co/datasets/hongpingzhang/rtx5090-energy-benchmark.rtx-5090-benchmarks
RTX 5090 LLM Benchmarks
Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig.
Quality Benchmarks
Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency.
Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/omegaprime669/rtx-5090-benchmarks.gr00t_rtx5090_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101",
"total_episodes": 60,
"total_frames": 17849,
"total_tasks": 1,
"total_videos": 60,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:60"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/akira-sasaki/gr00t_rtx5090_test.speculative-decoding-bench-rtx4090
Speculative Decoding Benchmark — RTX 4090
TL;DR: 4,576 benchmark runs measuring speculative decoding speedup / acceptance rate
across llama.cpp and LM Studio, Qwen3 (8B/14B) and Llama-3.1-8B target models, on a
single consumer RTX 4090 (24GB). Best observed case: the draft-free ngram-mod
self-speculative mode on structured tasks (JSON extraction 2.81x, code 2.76x,
global-median aggregation at temp=0). Open-ended tasks (creative writing, translation)
with a traditional draft… See the full description on the dataset page: https://huggingface.co/datasets/steven0226/speculative-decoding-bench-rtx4090.bon-high-reward-rtx4090
Best-of-N High-Reward Mode Coverage — RTX 4090 Experiment Results
Full experimental artifacts for the study of Best-of-N scaling, high-reward semantic mode
coverage, and antithetic-noise coupling for text-to-image diffusion (SDXL-Turbo, RTX 4090).
Experiments
#
Experiment
Pool
Key finding
1
Main BoN analysis — best-reward vs N, high-reward mode coverage (DINOv2 + agglomerative clustering), Vendi scores, fixed-N correlations, FE regression controlling HR… See the full description on the dataset page: https://huggingface.co/datasets/LamTNguyen/bon-high-reward-rtx4090.RTX3090-MoE-Cybersecurity-Evaluation
The 3090 Plan
A Seldon-inspired chronicle of seven model configurations, stubborn tools, and a statistically modest corner of the galaxy.
Updated: 2026-10-09T12:01:53Z. Hardware: 2 x RTX 3090 24 GiB, NVLink; 125 GiB system RAM. vLLM 0.31.0.
The Commission began with a reasonable proposition: a model that fails twice as quickly has not improved civilization. It has merely increased the rate at which somebody must intervene.
Seven serving configurations were brought before two RTX… See the full description on the dataset page: https://huggingface.co/datasets/Xananthium/RTX3090-MoE-Cybersecurity-Evaluation.qwen38-flash-next-int2-rtx5090-research
Qwen3.8-Flash-Next Mixed INT2 AutoRound on a Single RTX 5090
This benchmark and reproducibility artifact documents SGLang inference for the mixed-INT2 AutoRound Qwen3.8-Flash-Next checkpoint on one NVIDIA RTX 5090 Blackwell GPU. It covers a validated 256K / 262,144-token long context, MoE autotuning, tiered KV cache, CPU offload, and the device-local evidence showing why another Blackwell GPU's tuning configuration should not be copied blindly.
Headline inference… See the full description on the dataset page: https://huggingface.co/datasets/Jerrybro/qwen38-flash-next-int2-rtx5090-research.windows-rtx-4060ti-8gb-moe-offload-bench-2026-05
RTX 4060 Ti 8GB — Multi-Model Benchmark (2026-05)
practitioner benchmarks on consumer hardware (8GB VRAM, 32GB RAM). 10 models tested, covering MoE expert offload, hybrid SSM architectures, dense models, MLA, dense partial GPU offload, and the 1B speed ceiling. all runs on the same physical rig, same methodology.
current leaderboard (decode tok/s at sweet spot)
model
active params
GGUF size
sweet spot tok/s
quality (6 tests)
architecture
Llama 3.2 1B
1.24B
771… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/windows-rtx-4060ti-8gb-moe-offload-bench-2026-05.local-llms-benchmark-rtx5090
Local LLMs Benchmark — RTX 5090
Benchmark of 14 local language model configurations (9 distinct models, 27B–31B parameter range)
across 9 questions covering logical reasoning, Bayesian statistics, cognitive bias detection,
theoretical science, synthesis under contradiction, linguistic ambiguity, code optimization,
and AI ethics.
Hardware: RTX 5090 24GB | Intel Core Ultra 9 275HX | 64GB RAM | DebianInference backend: Ollama (Docker)Author: Francisco R. · LinkedIn
Full… See the full description on the dataset page: https://huggingface.co/datasets/Anodino/local-llms-benchmark-rtx5090.ubuntu22.04-rtx50series-blackwell-iso
Ubuntu 22.04 LTS 定制镜像 (专为 NVIDIA RTX 50 系列显卡优化)
📋 项目背景 (Project Background)
目前(2026年初),NVIDIA RTX 50 系列(Blackwell 架构,如 RTX 5080 / 5090)已正式发布。然而,由于硬件架构极新,传统的 Ubuntu 官方安装镜像在这些显卡上存在严重的兼容性问题。
本镜像由Jessy(杰西)制作,集成了最新的 NVIDIA 570 系列生产分支驱动,旨在为具身智能(Embodied AI)及深度学习开发者提供“开箱即用”的底层环境。
📋 为什么需要这个镜像?
如果你正在为新配的 RTX 50 系列(5080/5090 等) 工作站安装 Ubuntu 22.04,大概率会遇到死活进不去安装界面的情况。
🚫 传统的“救命药方”失效了
通常遇到显卡不兼容导致的黑屏,网上的常规解法是:
在 GRUB 界面开启 nomodeset(通用驱动模式)。
但实测证明: 对于架构大改的 RTX 50 系显卡,开启 nomodeset… See the full description on the dataset page: https://huggingface.co/datasets/Jessy-Huang/ubuntu22.04-rtx50series-blackwell-iso.test_rtx_3_cubeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 10,
"total_frames": 8862,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Askel1419/test_rtx_3_cube.bonsai2-27b-pq2-vs-fable-rtx5090
Bonsai 2 27B PQ2_0 vs Fable on a Single RTX 5090
This benchmark and reproducibility artifact compares single-GPU local LLM inference for Bonsai 2 27B / Qwen3.8 27B PQ2_0 GGUF through the Bonsai llama.cpp fork, its low-VRAM Q4_0 KV-cache quantization route with multimodal vision, and a Fable groupwise-int baseline. The controlled 210-question comparison on one RTX 5090 separates a quality-first Fable route from a 15,595 MiB sampled-peak Bonsai PQ2_0 + Q4_0-KV route.… See the full description on the dataset page: https://huggingface.co/datasets/Jerrybro/bonsai2-27b-pq2-vs-fable-rtx5090.test_rtx_pcThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 1,
"total_frames": 881,
"total_tasks": 1,
"total_videos": 4,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Askel1419/test_rtx_pc.chandra-ocr-2-rtx3090-vllm-benchmark
Chandra OCR 2 sur RTX 3090 - benchmark vLLM natif
Ce depot publie les resultats d'optimisation de datalab-to/chandra-ocr-2
sur une RTX 3090 24 Go RunPod, sans Docker ni sudo, avec le backend vLLM natif.
Resultat principal
configuration
pages/s
s/page
gain vs HF
qualite
HF origine
0.053
18.828
1.0x
stable
vLLM natif simple page
0.306
3.266
5.8x
stable
vLLM batch optimise
1.359
0.736
25.6x
stable
Le meilleur resultat observe est 1.359 pages/s… See the full description on the dataset page: https://huggingface.co/datasets/simonlesaumon/chandra-ocr-2-rtx3090-vllm-benchmark.test_rtx_3_cube_bboxesThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 10,
"total_frames": 8862,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/phospho-app/test_rtx_3_cube_bboxes.nanochat-rtx4070-sft-mixes
nanochat-rtx4070 SFT mixes
Eight SFT data mixes that were trained and evaluated on a single RTX 4070, and the results each one produced. Seven of them failed.
These are the actual independent variable behind the negative-results table in Bl4ckd09/nanochat-on-rtx4070. Every mix here was built deterministically, trained on the same frozen backbone with the same geometry and step count, and put through the same two-stage evaluation gate. Publishing only the winner would make the… See the full description on the dataset page: https://huggingface.co/datasets/Marcolini/nanochat-rtx4070-sft-mixes.LESA-FLUX-priority3-v3-RTX4090-20260927
INCOMPLETE ARCHIVE — NOT RESUMABLE
The 2026-09-27 RTX4090 campaign was not uploaded before the original VM became unavailable. This repository contains only surviving metadata in recovered_metadata_only: source snapshots, small locks and run logs. It has no training checkpoints, optimizer or RNG states, fresh teacher cache, conditioning tensors, or rollout vectors. It cannot support exact continuation from 2,000 updates. See the audit and manifest in that directory. No 5… See the full description on the dataset page: https://huggingface.co/datasets/LamTNguyen/LESA-FLUX-priority3-v3-RTX4090-20260927.200epPOKAchto_v2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/rtx409011/200epPOKAchto_v2.rtx5090-5m-bench
Manual RTX 5090 NVFP4 vLLM Benchmark Plan
Objective
Find the concurrency N that maximizes aggregate output token throughput on one rented RTX 5090 while every simulated user receives at least 15 output tokens/second over an exact five-minute measurement window.
The result is conditional on the fixed context-length mixture below. It is not a universal capacity number for every traffic distribution.
Primary optimization problem:
maximize aggregate_output_tok_s(N… See the full description on the dataset page: https://huggingface.co/datasets/suJayhh/rtx5090-5m-bench.ep_120This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/rtx409011/ep_120.octo-small-simpler-eggplant-50-rtx4500vast-rtx3090-market-6mo
Vast.ai RTX 3090 Spot Market, February-August 2026
Panel data from the vast.ai GPU rental marketplace, restricted to NVIDIA RTX 3090 offers. The public offer listing was polled every 10 minutes between 2026-02-13 and 2026-08-15. Each observation records price, hardware specifications, host reliability, and location. A derived lifecycle table gives the listing duration of every offer. Vast.ai does not publish historical listing data; this dataset was collected independently.… See the full description on the dataset page: https://huggingface.co/datasets/MarcusLammers/vast-rtx3090-market-6mo.octo-small-simpler-cube-stack-success-rtx4500
Octo-Small SIMPLER cube-stack success
This is a compact, auditable ManiSkill 3 evaluation artifact for
rail-berkeley/octo-small on SIMPLER's BridgeData-v2 digital twin task:
Environment: StackGreenCubeOnYellowCubeBakedTexInScene-v1
Instruction: stack the green block on the yellow block
Robot/policy setup: WidowX Bridge / Octo-Small
Hardware: NVIDIA RTX A4500, 20 GB
Date: 2026-07-22 UTC
Videos
Episode
Terminal result
Notes
48
Failure
Sustained grasp… See the full description on the dataset page: https://huggingface.co/datasets/lsnu/octo-small-simpler-cube-stack-success-rtx4500.rtx-5080-llm-throughput
RTX 5080 Local LLM Throughput: Measured, Not Estimated
Measured decode and prefill tokens/second, load times, and VRAM residency for
local LLMs on a single retail NVIDIA GeForce RTX 5080 (16 GB). 34 rows from two
dated captures:
2026-08-12, 16 rows: gpt-oss:20b (MXFP4), Qwen 2.5 14B and 7B (Q4_K_M)
and Llama 3.2 3B (Q4_K_M).
2026-08-26, 18 rows: the coding models qwen2.5-coder 7B and 14B,
qwen3-coder 30B and devstral 24B (all Q4_K_M), plus a gpt-oss:20b
re-capture for… See the full description on the dataset page: https://huggingface.co/datasets/iBlessi/rtx-5080-llm-throughput.200_epThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/rtx409011/200_ep.
