llmbenchio/benchmarks-by-vram
Updated on: 04 Oct 2026 Data contains: runs from the last 30 days Minimum runs: model/hardware combos with fewer than 3 runs are excluded llm-bench.io — Community LLM Benchmark Leaderboard by Hardware Per-model community benchmark data for local LLMs, curated from llm-bench.io and grouped by hardware and available VRAM. This dataset contains only aggregated statistics derived from individual benchmark submissions. It does not contain raw submissions, prompts, model responses… See the full description on the dataset page: https://huggingface.co/datasets/llmbenchio/benchmarks-by-vram.
<!-- AUTO-UPDATE:START --> Updated on: 04 Oct 2026 Data contains: runs from the last 30 days Minimum runs: model/hardware combos with fewer than 3 runs are excluded <!-- AUTO-UPDATE:END -->
llm-bench.io — Community LLM Benchmark Leaderboard by Hardware
Per-model community benchmark data for local LLMs, curated from llm-bench.io and grouped by hardware and available VRAM.
This dataset contains only aggregated statistics derived from individual benchmark submissions. It does not contain raw submissions, prompts, model responses, machine identifiers, or client/session data. The raw data lives in the llm-bench.io database and is used here only to compute the summary stats below.
Why this dataset
LLM enthusiasts want to know: "How well does this model run on my hardware, and how good is the output?" This table gives a trustworthy, community-sized answer by filtering to model/hardware combinations with enough independent runs to be meaningful.
What's in it
Methodology
Quality scores. Each submission includes scores on four tasks (code generation, long-horizon agent workflows, role-play, and research) that are scored by an automated LLM judge (a model used as an evaluator, not a human rater). Scores are 0–100.
Aggregation. For every (model, hardware) pair we:
- Take all runs from the last 30 days.
- Require at least 3 independent runs (this filters out noise and outliers — a single run is not trusted).
- Average the per-scenario scores across those runs.
Hardware grouping. Rows are grouped by the device a model was tested on, so a 32 tok/s figure on a laptop and a 32 tok/s figure on a desktop are never mixed.
Privacy. The published table is a statistical summary. Individual submissions, prompts, model outputs, machine hashes, and client/session identifiers are never exported — they remain only in the llm-bench.io database.
How to use
import pandas as pd
df = pd.read_csv("benchmarks-by-vram.csv")
# Best coding model on an RTX 4090 with at least 4 runs
df[(df.hardware.str.contains("4090")) & (df.run_count >= 4)] \
.sort_values("scenario_coding", ascending=False).head()Or load directly in Python:
from datasets import load_dataset
ds = load_dataset("llmbenchio/benchmarks-by-vram")License
Data is community-sourced and made available for reference and research use. Attribution is appreciated: source is [llm-bench.io](https://llm-bench.io).
