Team Ai
12 results

inference-benchmark

scaledown /vllm-inference-benchmarkstextn<1K0 likes1.1k downloads13d agoHugging Facehlarcher /inference-benchmarkertext100K<n<1M1 likes484 downloads2y agoHugging Facestudioburnside /mlx-local-inference-benchmarks MLX local-inference benchmarks — Qwen3.6 & Laguna-S/XS families Raw results, harnesses and methodology for an 8-axis benchmark of four MLX checkpoints on a 128 GB M5 Max. Everything a person would need to check my numbers or disagree with them. Companion model repos: Tess-4-27B-MLX-Q8 — with a working MTP head Tess-4-27B-MLX-Q4 — same, at 4-bit NEW (2026-07-24): the Laguna chapter — REPORT-LAGUNA.md + results-laguna/ Five-way same-engine bake-off (Laguna-S… See the full description on the dataset page: https://huggingface.co/datasets/studioburnside/mlx-local-inference-benchmarks.text-generation2 likes132 downloads3mo agoHugging Faceinference-optimization /speculators_benchmarks_tool_calltext1K<n<10K1 likes128 downloads6mo agoHugging Facehlarcher /inference-benchmarker-xettext100K<n<1M0 likes114 downloads2y agoHugging Facepxzleo /qwen3.8-27b-inference-benchmark-4090 Qwen3.8-27B Inference Benchmark on RTX 4090 48GB 中文说明 · GitHub benchmark repository Structured performance and accuracy results for four real Qwen3.8-27B serving configurations on an NVIDIA RTX 4090 48 GB workstation. A dual-GPU llama.cpp BF16 reference additionally used an RTX 3090 24 GB. This dataset is the analysis-friendly companion to the full benchmark repository. It publishes aggregate tables, 140 normalized per-request performance records, accuracy scores, sanitized… See the full description on the dataset page: https://huggingface.co/datasets/pxzleo/qwen3.8-27b-inference-benchmark-4090.tabularn<1K2 likes109 downloads2mo agoHugging Face