Team Ai
Datasetpublic

axjns/llmfit-benchmarks

llmfit Real-World LLM Inference Benchmarks An evolving dataset of real-world LLM inference measurements across consumer, workstation, datacenter, and unified-memory hardware. It combines benchmarks from an external community source with community benchmarks contributed directly to llmfit. The initial release contains 1,501 normalized observations: 1,010 unique external-community observations from the repository's 2026-08-10 snapshot. 491 llmfit-community observations from 61… See the full description on the dataset page: https://huggingface.co/datasets/axjns/llmfit-benchmarks.

sourceHugging Faceotherupdated 2mo agoView on Hugging Face
2likes128downloads
Dataset Card

llmfit Real-World LLM Inference Benchmarks

An evolving dataset of real-world LLM inference measurements across consumer, workstation, datacenter, and unified-memory hardware. It combines benchmarks from an external community source with community benchmarks contributed directly to llmfit.

The initial release contains 1,501 normalized observations:

  • —1,010 unique external-community observations from the repository's 2026-08-10 snapshot.
  • —491 llmfit-community observations from 61 validated submissions, including repository updates through 2026-08-19.

The external source reported 2,473 total measurements across 27 hardware presets at snapshot time, but the embedded llmfit cache retains at most 100 rows per preset. It contains 1,111 cached rows, of which 101 are exact duplicates appearing in overlapping hardware-preset buckets. This dataset publishes the 1,010 unique records; it does not represent the uncached records as available observations.

Sources and permission

Every row has a source value of external-community or llmfit-community, plus a stable source record identifier. The dataset maintainers received permission to redistribute the external benchmark data in this combined dataset; the source requested not to be named publicly. llmfit-community submissions are contributed through the llmfit repository and validated against its published schema.

This dataset uses license: other because it combines data from separately administered sources. The llmfit software repository is MIT licensed; that software license should not be interpreted as overriding rights attached to third-party model names, benchmark inputs, or source data.

Intended use

The dataset supports:

  • —Model-to-hardware fit analysis
  • —Throughput and latency estimation
  • —Estimate calibration against measured performance
  • —Runtime and quantization comparisons
  • —Hardware purchasing and deployment research
  • —Reproducibility studies across drivers and runtime versions

It is not a controlled leaderboard. Measurements come from heterogeneous machines, prompts, runtime configurations, software versions, thermals, and user environments.

Schema

The unified table preserves source provenance while normalizing common concepts:

  • —Identity: benchmark_id, source, source_record_id, measured_at
  • —Model: model_id, revision, family, parameter counts, MoE metadata
  • —Runtime: engine/provider, version, backend, quantization, engine options
  • —Hardware: accelerator, count, memory, CPU, OS, hardware class
  • —Workload: prompt/output tokens, context length, batch size, run count
  • —Results: output/prefill/total throughput, TTFT, duration, VRAM, and power
  • —Provenance: source snapshot, submitter name where already public, and verification status

Unavailable source fields are represented as null rather than inferred. Estimates and simulated results are not inserted as measured observations.

Updating

The export is deterministic. Run:

bash
python export.py /path/to/llmfit data/benchmarks.parquet

The llmfit community directory is append-only. Stable benchmark_id values make later exports suitable for deduplication and longitudinal use.

Limitations

  • —External-community data is a capped snapshot rather than a complete historical export.
  • —llmfit-community submissions currently contain aggregated runs rather than one row per raw run.
  • —Many records omit driver, runtime version, prompt, context, power, or peak-memory details.
  • —User-submitted measurements may be noisy or incorrectly identified.
  • —Results should not be treated as direct model-quality comparisons.

Citation

If this dataset is useful, cite the dataset revision you used and link to llmfit.