inference-benchmark
vllm-inference-benchmarksinference-benchmarkermlx-local-inference-benchmarks
MLX local-inference benchmarks — Qwen3.6 & Laguna-S/XS families
Raw results, harnesses and methodology for an 8-axis benchmark of four MLX
checkpoints on a 128 GB M5 Max. Everything a person would need to check my numbers
or disagree with them.
Companion model repos:
Tess-4-27B-MLX-Q8 — with a working MTP head
Tess-4-27B-MLX-Q4 — same, at 4-bit
NEW (2026-07-24): the Laguna chapter — REPORT-LAGUNA.md + results-laguna/
Five-way same-engine bake-off (Laguna-S… See the full description on the dataset page: https://huggingface.co/datasets/studioburnside/mlx-local-inference-benchmarks.speculators_benchmarks_tool_callinference-benchmarker-xetqwen3.8-27b-inference-benchmark-4090
Qwen3.8-27B Inference Benchmark on RTX 4090 48GB
中文说明 · GitHub benchmark repository
Structured performance and accuracy results for four real Qwen3.8-27B serving configurations on an NVIDIA RTX 4090 48 GB workstation. A dual-GPU llama.cpp BF16 reference additionally used an RTX 3090 24 GB.
This dataset is the analysis-friendly companion to the full benchmark repository. It publishes aggregate tables, 140 normalized per-request performance records, accuracy scores, sanitized… See the full description on the dataset page: https://huggingface.co/datasets/pxzleo/qwen3.8-27b-inference-benchmark-4090.
