Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01witcheer /rtx-5090-benchmarks RTX 5090 LLM Benchmarks Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig. Quality Benchmarks Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency. Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/rtx-5090-benchmarks.imagetext-generationn<1K1 likes1.5k downloads3d agoHugging Face02scaledown /vllm-inference-benchmarkstextn<1K0 likes1.1k downloads13d agoHugging Face03MothMalone /data-preprocessing-automl-benchmarks Data Preprocessing AutoML Benchmarks This repository contains text classification datasets with known data quality issues for preprocessing research in AutoML. Usage Load a specific dataset configuration like this: from datasets import load_dataset # Example for loading the TREC dataset dataset = load_dataset("MothMalone/data-preprocessing-automl-benchmarks", "trec") Available Datasets Below are the details for each dataset configuration available in this… See the full description on the dataset page: https://huggingface.co/datasets/MothMalone/data-preprocessing-automl-benchmarks.texttext-classification100K<n<1M0 likes476 downloads1y agoHugging Face04shubhxho /polymarket-hft-benchmarks Rust paper-engine benchmarks The local research desk exposes simulation playback, supplied-candidate ranking, benchmark provenance and retained local inference. Terminal GIF. Default operation has no paid API calls. Browser interaction/layout checks are unverified; no browser was connected and local TCP binding was refused by the validation sandbox. CPU API dispatch and packed-model inference were tested through the command-line transport. No company affiliation is claimed.… See the full description on the dataset page: https://huggingface.co/datasets/shubhxho/polymarket-hft-benchmarks.tabular10K<n<100K0 likes470 downloads15h agoHugging Face05dima0000 /ninfer-benchmarks NInfer on one RTX 5090 — recorded benchmark evidence Historical results from 19–20 September 2026, published by dima0000. This is a collection of benchmark evidence, not a model checkpoint or a live inference service. It preserves successful measurements, failed checks and incomplete experiments. Hardware: one NVIDIA RTX 5090 with 32 GB VRAM per trial. Models: regular and uncensored Qwen3.8-27B with NVFP4 weights. Engine: NInfer commit 9e163eee4b8acec21ab0ac765107b6a3f287b217… See the full description on the dataset page: https://huggingface.co/datasets/dima0000/ninfer-benchmarks.tabularn<1K0 likes312 downloads18d agoHugging Face06py-feat /benchmarks py-feat benchmarks Live benchmark data for py-feat and a cross-tool comparison against OpenFace 3.0, LibreFace, and PyAFAR. Powers the py-feat live dashboard. Updated by scheduled benchmark runs. Files File What accuracy.csv Tidy long table: one row per (tool, dataset, modality, metric). Covers AU F1 (DISFA+), 7-class emotion (AffectNet-val, RAF-DB), valence/arousal CCC (AffectNet-val), and gaze angular error (Columbia). throughput.csv py-feat… See the full description on the dataset page: https://huggingface.co/datasets/py-feat/benchmarks.tabularn<1K0 likes290 downloads4mo agoHugging Face07omegaprime669 /rtx-5090-benchmarks RTX 5090 LLM Benchmarks Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig. Quality Benchmarks Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency. Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/omegaprime669/rtx-5090-benchmarks.imagetext-generationn<1K0 likes218 downloads3mo agoHugging Face08diffusers /benchmarks Welcome to 🤗 Diffusers Benchmarks! This is dataset where we keep track of the inference latency and memory information of the core models in the diffusers library. Currently, the core models are: Flux Wan LTX SDXL Note that we will continue to extend this list based on their usage. You can analyze the results in this demo. [!IMPORTANT] Instead of benchmarking the entire diffusion pipelines, we only benchmark the forward passes of the diffusion networks under different settings… See the full description on the dataset page: https://huggingface.co/datasets/diffusers/benchmarks.tabularn<1K15 likes184 downloads9d agoHugging Face09mtapiapacheco /screen-benchmarkstext1M<n<10M0 likes163 downloads2mo agoHugging Face10llmbenchio /benchmarks-by-vramUpdated on: 04 Oct 2026 Data contains: runs from the last 30 days Minimum runs: model/hardware combos with fewer than 3 runs are excluded llm-bench.io — Community LLM Benchmark Leaderboard by Hardware Per-model community benchmark data for local LLMs, curated from llm-bench.io and grouped by hardware and available VRAM. This dataset contains only aggregated statistics derived from individual benchmark submissions. It does not contain raw submissions, prompts, model responses… See the full description on the dataset page: https://huggingface.co/datasets/llmbenchio/benchmarks-by-vram.tabularn<1K0 likes159 downloads6d agoHugging Face11humanlong /emotion-negotiation-benchmarks Emotion-Aware LLM Negotiation Benchmarks Four high-stakes, edge-deployable negotiation benchmarks — the official evaluation suite for our research program on emotion-aware LLM agents. Each benchmark targets a distinct domain where (a) LLM-vs-LLM negotiation has real-world consequences, and (b) on-device deployment of small language models matters for privacy and latency. The benchmarks were originally introduced with EmoMAS (ACL 2026 Main, top 9% of 12,148 submissions) and are… See the full description on the dataset page: https://huggingface.co/datasets/humanlong/emotion-negotiation-benchmarks.tabulartext-generationn<1K0 likes151 downloads4mo agoHugging Face12RaynarDM /apple-silicon-llm-benchmarks Apple Silicon Local LLM Benchmarks — M2 Max 32GB Measurements taken while trying to get Qwen3.8-27B usable locally on a 32GB M2 Max. Most of the popular speedup advice did not transfer from CUDA, so these are mostly negative results. Everything here was measured on one machine. Treat it as a datapoint, not a law. Hardware and software Chip Apple M2 Max Unified memory 32 GB (~21.8 GB wireable to the GPU) macOS 26.5.2 llama.cpp build c1d0e7a00… See the full description on the dataset page: https://huggingface.co/datasets/RaynarDM/apple-silicon-llm-benchmarks.tabularn<1K0 likes139 downloads12d agoHugging Face13sauravsingla08 /velographx-benchmarks VeloGraphX Benchmarks Machine-readable benchmark and reproducibility artifacts for VeloGraphX, a C++20 + Python engine for analytics on continuously evolving graphs. This Hugging Face repository is the benchmark/reproducibility companion to the canonical source repository at sauravsingla/VeloGraphX. It is published automatically from the GitHub main branch using a Hugging Face Trusted Publisher. The bundle is generated only from versioned repository artifacts; it does not invent… See the full description on the dataset page: https://huggingface.co/datasets/sauravsingla08/velographx-benchmarks.tabularn<1K0 likes130 downloads13d agoHugging Face14himajahealth /qwen3.8-27b-medical-benchmarks Qwen3.8-27B — Medical Benchmark Evaluation Evaluation report for Qwen/Qwen3.8-27B across 32 medical benchmark configurations (27,312 scored items), under a single fixed inference configuration, with two grading regimes: deterministic and model-graded. The model was evaluated as released. No weights were modified. Field Value Model under evaluation Qwen/Qwen3.8-27B Evaluation dates 2026-09-27 to 2026-09-30 (UTC) Benchmark configurations 32 (24 deterministic, 8… See the full description on the dataset page: https://huggingface.co/datasets/himajahealth/qwen3.8-27b-medical-benchmarks.tabularquestion-answeringn<1K0 likes129 downloads9d agoHugging Face15yuzhoucheng66 /HGBP-Benchmarks H-GBP Benchmark Inputs Exact numerical inputs for reproducing H-GBP and baseline experiments: 9 pose graphs and 10 bundle-adjustment problems. Each compressed file expands to the SHA-256-locked input used by the solver configurations. Download and run Code, build instructions and baseline adapters are available in the official Hierarchy-GBP GitHub repository. From the repository root, run: python scripts/download_datasets.py The downloader fetches only the… See the full description on the dataset page: https://huggingface.co/datasets/yuzhoucheng66/HGBP-Benchmarks.tabularn<1K1 likes119 downloads13d agoHugging Face16ruanjiange /whisper-browser-benchmarks whisper-browser-benchmarks Measurements from a Whisper transcription pipeline running entirely inside a browser tab: which audio and video containers the browser will actually decode, how accurate the smallest usable Whisper size is on clean synthetic speech, how long transcription takes relative to the length of the clip, what the first load pulls over the wire, and what happens to clips longer than the model's 30-second window. Everything here was measured, not quoted from a… See the full description on the dataset page: https://huggingface.co/datasets/ruanjiange/whisper-browser-benchmarks.audion<1K0 likes115 downloads24d agoHugging Face17manulife /tam-benchmarks Tasks over Application Manuals (TAM) TAM is a benchmark for evaluating long-horizon procedural reasoning: the ability of a language-model system to follow a large application manual, resolve cross-references, apply interdependent constraints, and produce an exact answer. Unlike short-horizon multi-hop tasks, TAM requires systems to maintain consistency across dozens of decisions drawn from manuals containing tens of thousands of rules. An early missed exception or incorrect… See the full description on the dataset page: https://huggingface.co/datasets/manulife/tam-benchmarks.documenttext-generationn<1K0 likes110 downloads29d agoHugging Face18genomic-benchmarks /genomic-benchmarks genomic-benchmarks, curated genomic-benchmarks, as published in quality-curated genomic benchmarks - one format, fixed row order, a permanent ID on every row. 7 datasets, 14 files, 600,434 rows, one gzipped CSV per split. Getting the data Two packages are the way in: genomic-benchmarks-data for people, genomic-benchmarks-data4agents for agents, the same functions either way. They resolve the URL, check the checksum, and carry each dataset's QC results, which this… See the full description on the dataset page: https://huggingface.co/datasets/genomic-benchmarks/genomic-benchmarks.tabular100K<n<1M0 likes105 downloads25d agoHugging Face19Agnuxo /optical-neuromorphic-eikonal-benchmarks Optical Neuromorphic Eikonal Solver - Benchmark Datasets Overview Benchmark datasets for evaluating the Optical Neuromorphic Eikonal Solver, a GPU-accelerated pathfinding algorithm achieving 30-300× speedup over CPU Dijkstra. 🎯 Key Results 134.9× average speedup vs CPU Dijkstra 0.64% mean error (sub-1% accuracy) 1.025× path length (near-optimal paths) 2-4ms per query on 512×512 grids 📊 Dataset Content 5 synthetic pathfinding test cases covering… See the full description on the dataset page: https://huggingface.co/datasets/Agnuxo/optical-neuromorphic-eikonal-benchmarks.tabularothern<1K0 likes86 downloads11mo agoHugging Face20genomic-benchmarks /miRBench miRBench, curated miRBench, as published in quality-curated genomic benchmarks - one format, fixed row order, a permanent ID on every row. 3 datasets, 6 files, 2,849,872 rows, one gzipped CSV per split, all Homo sapiens. Getting the data Two packages are the way in: genomic-benchmarks-data for people, genomic-benchmarks-data4agents for agents, the same functions either way. They resolve the URL, check the checksum, and carry each dataset's QC results, which this… See the full description on the dataset page: https://huggingface.co/datasets/genomic-benchmarks/miRBench.tabular1M<n<10M0 likes74 downloads25d agoHugging Face21derekl35 /quantization-benchmarkstabularn<1K3 likes71 downloads1y agoHugging Face22Djangodevreng /dgx-spark-benchmarks DGX Spark LLM Arena benchmarks Reproducible LLM inference benchmarks on an NVIDIA DGX Spark (GB10, 128 GB unified memory). The suite defines eleven tests: six closed-loop (llama-benchy) and five open-loop (vllm bench serve). Results cover all eleven: the ten throughput tests under results, and the rate sweep under rateSweep. Raw results remain inspectable, but only complete runs without a failed sanity check count toward rankings and aggregate throughput. Open-loop tests must… See the full description on the dataset page: https://huggingface.co/datasets/Djangodevreng/dgx-spark-benchmarks.tabularn<1K1 likes66 downloads22d agoHugging Face23Maverick03511 /prepqc-exploitation-benchmarks PrePQC Exploitation: public benchmark observations Four separate small experiment records from the PrePQC Exploitation source repository, run on 2026-10-08. Each row is an aggregate of the stated job, operation, or fixed-seed trials; rows are not raw per-shot/per-case records or independent replications. These tables support reproduction and plotting; they are not a training set, a leaderboard, evidence of quantum speedup, or evidence that production ML-KEM was broken. The… See the full description on the dataset page: https://huggingface.co/datasets/Maverick03511/prepqc-exploitation-benchmarks.tabularn<1K0 likes62 downloads1d agoHugging Face24gamlin /deepseek-v41-flash-h200-vs-opus-benchmarks DeepSeek vs Opus: The Numbers Don't Add Up The measured numbers behind DeepSeek vs Opus: The Numbers Don't Add Up by The Call Center Doctors. Shared with credit and a link back, as the original allows. The original has the full fact sheet. We rented 4 H200s to run DeepSeek V4.1 Flash (763B parameters, 1M-token context) on vLLM, pointed Claude Code at it, and compared it with Opus 5.5 on real coding work. All numbers were measured on 2026-09-27 and from our own September (Sep… See the full description on the dataset page: https://huggingface.co/datasets/gamlin/deepseek-v41-flash-h200-vs-opus-benchmarks.textn<1K0 likes60 downloads13d agoHugging Face25ValorSME /sme-valuation-benchmarks-2026 SME Valuation Benchmarks 2026 Reference dataset for small and medium-sized enterprise (SME) valuation: discount rates (WACC), unlevered sector betas and EV/EBITDA multiple ranges for 11 industry sectors across 12 countries (France, Spain, Germany, United Kingdom, United States, Australia, Singapore, India, New Zealand, Ireland, Canada, South Africa). 110 rows. Columns Column Description sector Sector key (e.g. software-saas, construction) sector_label… See the full description on the dataset page: https://huggingface.co/datasets/ValorSME/sme-valuation-benchmarks-2026.tabularn<1K0 likes57 downloads29d agoHugging Face26mtapiapacheco /GUE-benchmarkstext100K<n<1M0 likes53 downloads2mo agoHugging Face27llmspeed /llm-speed-benchmarks llm-speed: signed LLM inference-speed benchmarks Crowdsourced, cryptographically signed measurements of how fast large language models actually run: decode tokens per second, time to first token, and latency, across consumer GPUs, Apple Silicon, and hosted APIs, under one reproducible workload suite (suite-v1). Live data and bulk downloads: https://llm-speed.com/data Per-run permalink: https://llm-speed.com/r/<id> Methodology: https://llm-speed.com/methodology DOI:… See the full description on the dataset page: https://huggingface.co/datasets/llmspeed/llm-speed-benchmarks.textn<1K0 likes51 downloads3mo agoHugging Face28ashwin-sreedhar /runpod-serverless-benchmarks Runpod Serverless LLM benchmarks Raw benchmark output from running LLM inference on Runpod Serverless between 2026-09-23 and 2026-10-01: cold-start timelines, load sweeps, per-request latencies, the workers' own boot logs, and an LLM scoring job run against several backends. Two engines appear on the same GPU and model — emberserve (my from-scratch inference engine, formerly pagedserve) and Runpod's worker-vllm — which is what makes the data worth keeping. This is the landing… See the full description on the dataset page: https://huggingface.co/datasets/ashwin-sreedhar/runpod-serverless-benchmarks.tabular1K<n<10K0 likes48 downloads7d agoHugging Face29docs-benchmarks /compile-benchmarkstabularn<1K0 likes43 downloads2y agoHugging Face30EvoLenTokenizer /ATAC-benchmarkstext100K<n<1M0 likes36 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.