datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vllm-inference-benchmarksinference-benchmarkermlx-local-inference-benchmarks
MLX local-inference benchmarks — Qwen3.6 & Laguna-S/XS families
Raw results, harnesses and methodology for an 8-axis benchmark of four MLX
checkpoints on a 128 GB M5 Max. Everything a person would need to check my numbers
or disagree with them.
Companion model repos:
Tess-4-27B-MLX-Q8 — with a working MTP head
Tess-4-27B-MLX-Q4 — same, at 4-bit
NEW (2026-07-24): the Laguna chapter — REPORT-LAGUNA.md + results-laguna/
Five-way same-engine bake-off (Laguna-S… See the full description on the dataset page: https://huggingface.co/datasets/studioburnside/mlx-local-inference-benchmarks.speculators_benchmarks_tool_callinference-benchmarker-xetqwen3.8-27b-inference-benchmark-4090
Qwen3.8-27B Inference Benchmark on RTX 4090 48GB
中文说明 · GitHub benchmark repository
Structured performance and accuracy results for four real Qwen3.8-27B serving configurations on an NVIDIA RTX 4090 48 GB workstation. A dual-GPU llama.cpp BF16 reference additionally used an RTX 3090 24 GB.
This dataset is the analysis-friendly companion to the full benchmark repository. It publishes aggregate tables, 140 normalized per-request performance records, accuracy scores, sanitized… See the full description on the dataset page: https://huggingface.co/datasets/pxzleo/qwen3.8-27b-inference-benchmark-4090.hf-inference-endpoint-benchmarks
Raw benchmark result files
Raw JSON outputs from the sessions described in benchmarking-methodology.md. Model: Qwen3.5-4B family, hf-endpoints deployed via the configs documented in cli-and-api.md.
Short-prompt decode comparison (valid metric — prompt negligible vs output, see trap #1 in methodology doc)
File
Setup
llamacpp_results.json
llama.cpp, GGUF Q8_0, MTP, A10G
vllm_results.json
vLLM, FP8-dynamic, MTP, A10G
vllm_bf16_results.json
vLLM, bf16… See the full description on the dataset page: https://huggingface.co/datasets/LostGentoo/hf-inference-endpoint-benchmarks.inference-benchmarkerenterprise-llm-inference-benchmarks-2026
🚀 Enterprise LLM Inference & Fine-Tuning Benchmarks (2026 Guide)
A curated benchmark index and architectural guide evaluating open-source foundation models, real-time inference engines (vLLM vs. TensorRT-LLM), and cloud GPU economics for enterprise deployments.
🧠 Open-Source Foundation Model Benchmarks (RAG & Code Generation)
Flagship Evaluation: Top Open-Source LLMs for Enterprise RAG & Code Generation (2026 In-Depth Guide) — Comparing Qwen 2.5 Coder, Llama… See the full description on the dataset page: https://huggingface.co/datasets/Abdulrahmankalil/enterprise-llm-inference-benchmarks-2026.inference-video-benchmark-v1
Inference Video Benchmark v1
This directory contains the video benchmark dataset used for local Gemma 4 multimodal serving evaluation.
Contents
clips/: extracted benchmark clips
clips.jsonl: clip manifest
sources.jsonl: source-video manifest with attribution and license fields
summary.json: dataset counts and metadata
Local-only artifacts:
sources/: downloaded full source videos
clips.local.jsonl: local clip manifest with machine-local paths
Dataset Shape… See the full description on the dataset page: https://huggingface.co/datasets/younes-ovs/inference-video-benchmark-v1.edge-inference-benchmarks
TinyEdge edge-inference benchmarks
Independently measured latency and accuracy for well-known vision models on
real edge devices (phones, tablets — fleet growing), produced by
TinyEdge, a device cloud for edge-AI benchmarking.
Nothing here is taken from papers or spec sheets: every row is a job executed
on the physical device through TinyEdge's production agent, with accuracy
measured on a fixed 500-image stratified sample of
ImageNet-V2 (matched-frequency)
using a standardized… See the full description on the dataset page: https://huggingface.co/datasets/TinyEdge/edge-inference-benchmarks.timm_inference_benchmarkllm-inference-benchmarkinference-text-benchmark-v1
Inference Text Benchmark v1
This directory contains the text benchmark dataset used for local Gemma 4 serving and throughput evaluation.
Contents
prompts.jsonl: final benchmark prompts
sources.curated.jsonl: source-segment manifest with provenance
summary.json: dataset counts and metadata
Dataset Shape
300 prompts total
50 prompts in each target input bucket:
128
512
2048
8192
32768
65536
Each prompt record includes:
Stable prompt_id
bucket_target_tokens… See the full description on the dataset page: https://huggingface.co/datasets/younes-ovs/inference-text-benchmark-v1.moe-inference-benchmark
Systematic Architecture Search for Mobile-Optimized Mixture of Experts Language Models
Authors: Kshitij Thakkar
Date: February 2026
Collection: Mobile MoE Architecture Search (32 models)
Dataset: kshitijthakkar/moe-inference-benchmark
Abstract
We present a systematic architecture search for Mixture of Experts (MoE) language models optimized for mobile deployment via GGUF quantization. Through 41 experiments exploring model size, expert count, routing strategies… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/moe-inference-benchmark.large-moe-inference-benchmark
