rtx
Qwen3.8-27B-NVFP4-RTX5090Qwen3.8-27B-Uncensored-NVFP4-RTX5090Qwen3.8-27B-NVFP4-RTX5090-LMHead4Qwen3.8-27B-RTX5080Qwen3.8-27B-Uncensored-DSpark-RTX5090Swift-1.5-Qwen3.8-27B-W4A16-RTX3090-NInfer-v3Huihui-Qwen3.8-27B-Abliterated-Gittensor-Style-NVFP4-RTX5090DeepSeek-V4.1-Flash-EXL3-3bpw-2x-RTX-PRO-6000
Datasets
All datasets matching “rtx”LESA-FLUX-A100-cutpoint-RTX4090-20260928
A100-origin cutpoint follow-up
This is a new research campaign from the pinned public A100 artifact revision 45671fe9ea2bd97491d921338cbf85a853e34578, not a continuation of the lost fresh-RTX4090 2k checkpoints. The origin teacher cache and base model are referenced by revision/SHA rather than copied here. New measurements and any new checkpoints/derived tensors will be uploaded by stage. Until verified result folders appear, this repository contains only the plan, not completed… See the full description on the dataset page: https://huggingface.co/datasets/LamTNguyen/LESA-FLUX-A100-cutpoint-RTX4090-20260928.GenEmotions-RTX5070Tirtx-5090-benchmarks
RTX 5090 LLM Benchmarks
Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig.
Quality Benchmarks
Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency.
Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/rtx-5090-benchmarks.multigpu-beam-rtx3060-scaling-20261004
Completed four-series scaling measurement
Full report · Machine-readable results · CSV table · Article fragment · Experiment protocol
We measured MultiGPU Beam Search on a single host with eight identical RTX 3060 12 GiB GPUs, using 1, 2, 4 and 8 devices. The fixed-work experiment used a global beam of 4,194,304 and twelve complete search steps, expanding 637,616,736 actions in every run. With the common execution profile, median synchronized end-to-end search times were 85.107… See the full description on the dataset page: https://huggingface.co/datasets/TryDotAtwo/multigpu-beam-rtx3060-scaling-20261004.qwen3-coder-gb10-vs-rtx5090-benchmark
NVIDIA GB10 vs. GeForce RTX 5090 - Local LLM Inference Benchmark
Model: Qwen3-Coder-30B-A3B-InstructFormat: GGUF, Q4_K_M, 18.63 GBRuntime: LM Studio / llama.cppAuthor: Efehan A.Benchmark date: 5 August 2026
This repository contains a decode-focused local inference benchmark comparing an NVIDIA GB10 system with a Windows workstation containing two GeForce RTX 5090 GPUs. Telemetry shows that the inference workload was carried primarily by a single RTX 5090 (GPU 0), while GPU 1… See the full description on the dataset page: https://huggingface.co/datasets/mreltera/qwen3-coder-gb10-vs-rtx5090-benchmark.rtx5090-energy-benchmark
RTX 5090 LLM Energy Benchmark
First energy efficiency benchmark of 4-bit quantization on NVIDIA RTX 5090 (Blackwell architecture).
Key Finding
4-bit quantization increases energy consumption by up to 29% for models < 5B parameters.
The crossover point where quantization becomes beneficial is ~5B parameters.
Results
Model
FP16 Energy
4-bit Energy
Change
TinyLlama 1.1B
1,659 J/1k
2,098 J/1k
+26.5% 🔴
Qwen2 1.5B
2,411 J/1k
3,120 J/1k
+29.4% 🔴… See the full description on the dataset page: https://huggingface.co/datasets/hongpingzhang/rtx5090-energy-benchmark.
