datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mlx-local-inference-benchmarks
MLX local-inference benchmarks — Qwen3.6 & Laguna-S/XS families
Raw results, harnesses and methodology for an 8-axis benchmark of four MLX
checkpoints on a 128 GB M5 Max. Everything a person would need to check my numbers
or disagree with them.
Companion model repos:
Tess-4-27B-MLX-Q8 — with a working MTP head
Tess-4-27B-MLX-Q4 — same, at 4-bit
NEW (2026-07-24): the Laguna chapter — REPORT-LAGUNA.md + results-laguna/
Five-way same-engine bake-off (Laguna-S… See the full description on the dataset page: https://huggingface.co/datasets/studioburnside/mlx-local-inference-benchmarks.hf-inference-endpoint-benchmarks
Raw benchmark result files
Raw JSON outputs from the sessions described in benchmarking-methodology.md. Model: Qwen3.5-4B family, hf-endpoints deployed via the configs documented in cli-and-api.md.
Short-prompt decode comparison (valid metric — prompt negligible vs output, see trap #1 in methodology doc)
File
Setup
llamacpp_results.json
llama.cpp, GGUF Q8_0, MTP, A10G
vllm_results.json
vLLM, FP8-dynamic, MTP, A10G
vllm_bf16_results.json
vLLM, bf16… See the full description on the dataset page: https://huggingface.co/datasets/LostGentoo/hf-inference-endpoint-benchmarks.enterprise-llm-inference-benchmarks-2026
🚀 Enterprise LLM Inference & Fine-Tuning Benchmarks (2026 Guide)
A curated benchmark index and architectural guide evaluating open-source foundation models, real-time inference engines (vLLM vs. TensorRT-LLM), and cloud GPU economics for enterprise deployments.
🧠 Open-Source Foundation Model Benchmarks (RAG & Code Generation)
Flagship Evaluation: Top Open-Source LLMs for Enterprise RAG & Code Generation (2026 In-Depth Guide) — Comparing Qwen 2.5 Coder, Llama… See the full description on the dataset page: https://huggingface.co/datasets/Abdulrahmankalil/enterprise-llm-inference-benchmarks-2026.moe-inference-benchmark
Systematic Architecture Search for Mobile-Optimized Mixture of Experts Language Models
Authors: Kshitij Thakkar
Date: February 2026
Collection: Mobile MoE Architecture Search (32 models)
Dataset: kshitijthakkar/moe-inference-benchmark
Abstract
We present a systematic architecture search for Mixture of Experts (MoE) language models optimized for mobile deployment via GGUF quantization. Through 41 experiments exploring model size, expert count, routing strategies… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/moe-inference-benchmark.
