Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01witcheer /rtx-5090-benchmarks RTX 5090 LLM Benchmarks Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig. Quality Benchmarks Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency. Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/rtx-5090-benchmarks.imagetext-generationn<1K1 likes1.5k downloads3d agoHugging Face02omegaprime669 /rtx-5090-benchmarks RTX 5090 LLM Benchmarks Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig. Quality Benchmarks Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency. Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/omegaprime669/rtx-5090-benchmarks.imagetext-generationn<1K0 likes218 downloads3mo agoHugging Face03yshenaw /SkillOpt_Lite_Benchmarks SkillOpt_Lite Benchmarks Train / val / test splits used by the SkillOpt_Lite project. One multi-config repo containing all six benchmarks: Config Rows (train / val / test) Content shipped searchqa 400 / 200 / 1400 Full QA — id, question, list of DOC contexts, answers. Sampled from dl4ir-searchQA. docvqa 107 / 53 / 374 Full QA + images bundled — parquet has id/question/answers/topic/image_path; PNGs live under docvqa_images/ at the repo root. Subset of… See the full description on the dataset page: https://huggingface.co/datasets/yshenaw/SkillOpt_Lite_Benchmarks.imagequestion-answering1K<n<10K0 likes213 downloads3mo agoHugging Face04LemOneLabs /OMNIX_Benchmarks_Latest 📊 OMNIX Benchmark Comparison Table Model Overall Score (Grade) Format Adherence Logical Reasoning Knowledge Recall Constraint Following First-Pass SR Eventual SR FCI Avg Latency (ms) qwen-3-4b-q4 92/100 (A) 97/100 77/100 100/100 95/100 89% 95% 0.56 16140 gemma-4-e4b-q4 89/100 (B) 100/100 67/100 100/100 94/100 94% 94% 0.22 34308 qwen-2.5-coder-3b-text 76/100 (C) 94/100 50/100 95/100 68/100 72% 87% 1.11 3601 llama-3.2-3b-q4 75/100 (C) 89/100 37/100 98/100 84/100… See the full description on the dataset page: https://huggingface.co/datasets/LemOneLabs/OMNIX_Benchmarks_Latest.documenttext-generationn<1K0 likes207 downloads3mo agoHugging Face05ChuGyouk /Qwen3.5-4B-nothink-benchmarks2026-09-27 update: RealMath benchmark results added. Qwen3.5-4B (non-thinking) — 14 benchmarks, multi-sample outputs with pass@k All sampled outputs of Qwen/Qwen3.5-4B in non-thinking mode (enable_thinking=False) on 14 benchmarks, with per-response correctness and pass@k / avg@n metrics. One row per problem; every row carries the exact prompt that was sent to the model, all sampled responses, their scores, and the benchmark-level metrics. The 2026-09-27 update adds all 1,286… See the full description on the dataset page: https://huggingface.co/datasets/ChuGyouk/Qwen3.5-4B-nothink-benchmarks.texttext-generation10K<n<100K0 likes198 downloads13d agoHugging Face06noel7Y /data-agent-benchmarks LongHorizon Full Data-Agent Benchmarks Companion data artifacts for five complete evaluation tracks: DataSciBench full55 / 167 metric entries DABStep full450 DABStep-Research full100 DSBench Modeling full74 LongDS full68 / 2,225 turns The companion GitHub repository contains processed manifests, evaluation code, historical API ReAct baseline code, download/preparation tools, and the frozen source lock. artifact_manifest.json records every uploaded object's size, SHA-256… See the full description on the dataset page: https://huggingface.co/datasets/noel7Y/data-agent-benchmarks.imagequestion-answering1 likes178 downloads1mo agoHugging Face07humanlong /emotion-negotiation-benchmarks Emotion-Aware LLM Negotiation Benchmarks Four high-stakes, edge-deployable negotiation benchmarks — the official evaluation suite for our research program on emotion-aware LLM agents. Each benchmark targets a distinct domain where (a) LLM-vs-LLM negotiation has real-world consequences, and (b) on-device deployment of small language models matters for privacy and latency. The benchmarks were originally introduced with EmoMAS (ACL 2026 Main, top 9% of 12,148 submissions) and are… See the full description on the dataset page: https://huggingface.co/datasets/humanlong/emotion-negotiation-benchmarks.tabulartext-generationn<1K0 likes151 downloads4mo agoHugging Face08CopyleftCultivars /qwen3.6-35b-a3b-chemistry-benchmarks Qwen3.6-35B-A3B Chemistry Benchmark Results Raw outputs and scores from running Qwen3.6-35B-A3B (Q8_0 quant) through five published chemistry and biosecurity benchmarks, entirely on local hardware (two secondhand Tesla M40 24GB GPUs, no cloud compute). This is the raw data behind our blog post on locally reproducible AI capability evaluation, including the full per-item outputs, the parsing failures, and the negative results, not just the headline numbers. Results at… See the full description on the dataset page: https://huggingface.co/datasets/CopyleftCultivars/qwen3.6-35b-a3b-chemistry-benchmarks.question-answering0 likes145 downloads2mo agoHugging Face09nyamtulla /benchmarking-the-benchmarks-data Benchmarking the Benchmarks — Raw SLM Safety Evaluation Runs Raw evaluation data for the ESORICS 2026 paper: Nyamtulla Shaik, Fengjun Li, Bo Luo — University of Kansas Benchmarking the Benchmarks: Evaluating Automated Safety Benchmarks for Small Language Models. ESORICS 2026. arXiv:2608.17183 Code and processed data: https://github.com/nyamtulla/benchmarking-the-benchmarks ⚠️ Content warning This dataset contains adversarial safety prompts and model responses… See the full description on the dataset page: https://huggingface.co/datasets/nyamtulla/benchmarking-the-benchmarks-data.text-generation100K<n<1M2 likes144 downloads1mo agoHugging Face10studioburnside /mlx-local-inference-benchmarks MLX local-inference benchmarks — Qwen3.6 & Laguna-S/XS families Raw results, harnesses and methodology for an 8-axis benchmark of four MLX checkpoints on a 128 GB M5 Max. Everything a person would need to check my numbers or disagree with them. Companion model repos: Tess-4-27B-MLX-Q8 — with a working MTP head Tess-4-27B-MLX-Q4 — same, at 4-bit NEW (2026-07-24): the Laguna chapter — REPORT-LAGUNA.md + results-laguna/ Five-way same-engine bake-off (Laguna-S… See the full description on the dataset page: https://huggingface.co/datasets/studioburnside/mlx-local-inference-benchmarks.text-generation2 likes132 downloads3mo agoHugging Face11himajahealth /qwen3.8-27b-medical-benchmarks Qwen3.8-27B — Medical Benchmark Evaluation Evaluation report for Qwen/Qwen3.8-27B across 32 medical benchmark configurations (27,312 scored items), under a single fixed inference configuration, with two grading regimes: deterministic and model-graded. The model was evaluated as released. No weights were modified. Field Value Model under evaluation Qwen/Qwen3.8-27B Evaluation dates 2026-09-27 to 2026-09-30 (UTC) Benchmark configurations 32 (24 deterministic, 8… See the full description on the dataset page: https://huggingface.co/datasets/himajahealth/qwen3.8-27b-medical-benchmarks.tabularquestion-answeringn<1K0 likes129 downloads9d agoHugging Face12axjns /llmfit-benchmarks llmfit Real-World LLM Inference Benchmarks An evolving dataset of real-world LLM inference measurements across consumer, workstation, datacenter, and unified-memory hardware. It combines benchmarks from an external community source with community benchmarks contributed directly to llmfit. The initial release contains 1,501 normalized observations: 1,010 unique external-community observations from the repository's 2026-08-10 snapshot. 491 llmfit-community observations from 61… See the full description on the dataset page: https://huggingface.co/datasets/axjns/llmfit-benchmarks.tabulartext-generation1K<n<10K2 likes128 downloads2mo agoHugging Face13RISys-Lab /Benchmarks_CyberSec_SECURE Dataset Card for SECURE (RISys-Lab Mirror) ⚠️ Disclaimer: > This repository is a mirror/re-host of the original SECURE benchmark.RISys-Lab is not the author of this dataset. We are hosting this copy in Parquet format to ensure seamless integration and stability for our internal evaluation pipelines. All credit belongs to the original authors listed below. Repository Intent This Hugging Face dataset is a re-host of the original SECURE benchmark. It has been converted… See the full description on the dataset page: https://huggingface.co/datasets/RISys-Lab/Benchmarks_CyberSec_SECURE.texttext-classification1K<n<10K0 likes112 downloads9mo agoHugging Face14callensxavier /runux-tpu-v5e-benchmarks ⚡ RunuX-AI — TPU v5e Inference Benchmarks Achieving 3× Throughput & 3× Energy Reduction on Google TPU v5e Xavier Callens · Socrate AI Lab (Non-Profit) Reproducible benchmark data & scripts — No proprietary code included 🎯 What Is This? This repository contains benchmark results and Apache-2.0 reproduction scripts for comparing LLM inference performance across 5 frameworks on Google TPU v5e. The goal is to enable independent verification of our claims… See the full description on the dataset page: https://huggingface.co/datasets/callensxavier/runux-tpu-v5e-benchmarks.texttext-generationn<1K0 likes111 downloads5mo agoHugging Face15manulife /tam-benchmarks Tasks over Application Manuals (TAM) TAM is a benchmark for evaluating long-horizon procedural reasoning: the ability of a language-model system to follow a large application manual, resolve cross-references, apply interdependent constraints, and produce an exact answer. Unlike short-horizon multi-hop tasks, TAM requires systems to maintain consistency across dozens of decisions drawn from manuals containing tens of thousands of rules. An early missed exception or incorrect… See the full description on the dataset page: https://huggingface.co/datasets/manulife/tam-benchmarks.documenttext-generationn<1K0 likes110 downloads29d agoHugging Face16LostGentoo /hf-inference-endpoint-benchmarks Raw benchmark result files Raw JSON outputs from the sessions described in benchmarking-methodology.md. Model: Qwen3.5-4B family, hf-endpoints deployed via the configs documented in cli-and-api.md. Short-prompt decode comparison (valid metric — prompt negligible vs output, see trap #1 in methodology doc) File Setup llamacpp_results.json llama.cpp, GGUF Q8_0, MTP, A10G vllm_results.json vLLM, FP8-dynamic, MTP, A10G vllm_bf16_results.json vLLM, bf16… See the full description on the dataset page: https://huggingface.co/datasets/LostGentoo/hf-inference-endpoint-benchmarks.text-generationn<1K1 likes95 downloads3mo agoHugging Face17RLAIF-V /RLPR-Benchmarks Dataset Card for RLPR-Test GitHub | Paper News: [2025.06.23] 📃 Our paper detailing the RLPR framework and its comprehensive evaluation using this suite is accessible at arXiv! Dataset Summary We include the following seven benchmarks for evaluation of RLPR: Mathematical Reasoning Benchmarks: MATH-500 (Cobbe et al., 2021) Minerva (Lewkowycz et al., 2022) AIME24 General Domain Reasoning Benchmarks: MMLU-Pro (Wang et al., 2024): A multitask language… See the full description on the dataset page: https://huggingface.co/datasets/RLAIF-V/RLPR-Benchmarks.textquestion-answeringn<1K0 likes90 downloads1y agoHugging Face18Abdulrahmankalil /enterprise-llm-inference-benchmarks-2026 🚀 Enterprise LLM Inference & Fine-Tuning Benchmarks (2026 Guide) A curated benchmark index and architectural guide evaluating open-source foundation models, real-time inference engines (vLLM vs. TensorRT-LLM), and cloud GPU economics for enterprise deployments. 🧠 Open-Source Foundation Model Benchmarks (RAG & Code Generation) Flagship Evaluation: Top Open-Source LLMs for Enterprise RAG & Code Generation (2026 In-Depth Guide) — Comparing Qwen 2.5 Coder, Llama… See the full description on the dataset page: https://huggingface.co/datasets/Abdulrahmankalil/enterprise-llm-inference-benchmarks-2026.tabulartext-generationn<1K1 likes80 downloads23d agoHugging Face19zevatov /nra-benchmarks 🧬 NRA Benchmark Datasets All benchmark datasets for Neural Ready Archive (NRA) — the Rust-native streaming format for ML training. Train on gigabytes of real data without downloading a single byte. NRA replaces tar.gz and zip for the AI era. 📦 Available Datasets File Domain Source Files Size food-101.nra 🖼️ Vision ethz/food101 101,000 images 4.7 GB wikitext.nra 📝 Text Salesforce/wikitext 23,767 text files 7.6 MB pokemon.nra 🎨 Multimodal… See the full description on the dataset page: https://huggingface.co/datasets/zevatov/nra-benchmarks.image-classification100K<n<1M1 likes71 downloads5mo agoHugging Face20LemOneLabs /OMNIX_Benchmarks_6-30-26 📊 OMNIX Benchmark Comparison Table Model Overall Eventual Score (Grade) Format Adherence (Eventual) Logical Reasoning (Eventual) Knowledge Recall (Eventual) Constraint Following (Eventual) First-Pass Success Rate Eventual Success Rate Friction Correction Index (FCI) Average Request Latency gemma-3 1B 65/100 (D) 97/100 23/100 90/100 59/100 60% 74% 1.56 9545ms gemma-4-e2b-q4 73/100 (C) 100/100 33/100 88/100 83/100 72% 83% 0.78 20575ms gemma-4-e4b-q4 89/100 (B)… See the full description on the dataset page: https://huggingface.co/datasets/LemOneLabs/OMNIX_Benchmarks_6-30-26.documenttext-generationn<1K0 likes61 downloads3mo agoHugging Face21zhangdw /Anchor-benchmarks 🧠 Anchor Benchmarks A curated long-term memory benchmark bundle for LLM and agent evaluation &nbsp;&nbsp;&nbsp;&nbsp; Anchor Benchmarks packages three public long-term memory evaluation resources for studying factual recall, temporal reasoning, knowledge update, multi-hop inference, and multimodal conversational memory. Quick Start · At a Glance · Benchmarks · Evaluation · Citation [!IMPORTANT] This repository is a benchmark bundle, not a new… See the full description on the dataset page: https://huggingface.co/datasets/zhangdw/Anchor-benchmarks.question-answering1K<n<10K0 likes59 downloads4mo agoHugging Face22hmnshudhmn24 /llm-benchmarks-capabilities-2020-2026 📊 LLM Benchmarks & Capabilities 2020–2026 The most comprehensive open dataset tracking the evolution of Large Language Models — from GPT-3 to GPT-5.5, Claude Opus 4.7, Gemini 3.5, and beyond. 🧭 Overview This dataset captures the complete LLM landscape from 2020 to 2026 across five dimensions: 🤖 113 models from 25+ organizations 📈 17 benchmarks tracking capability growth over time 💰 Monthly API pricing showing 100x+ cost reductions ⚙️ Training compute… See the full description on the dataset page: https://huggingface.co/datasets/hmnshudhmn24/llm-benchmarks-capabilities-2020-2026.tabulartext-generation1 likes55 downloads4mo agoHugging Face23LemOneLabs /OMNIX_Benchmarks_7-6-26 📊 OMNIX Benchmark Comparison Table Model Overall Score (Grade) Format Adherence Logical Reasoning Knowledge Recall Constraint Following First-Pass SR Eventual SR FCI Avg Latency (ms) qwen-3-4b-q4 92/100 (A) 97/100 77/100 100/100 95/100 89% 95% 0.56 16140 llama-3.2-3b-q4 75/100 (C) 89/100 37/100 98/100 84/100 78% 84% 0.89 5559 bonsai-8b-q4 63/100 (D) 89/100 13/100 88/100 76/100 75% 76% 1.11 9721 Note: SR = Success Rate. FCI (Friction Correction Index)… See the full description on the dataset page: https://huggingface.co/datasets/LemOneLabs/OMNIX_Benchmarks_7-6-26.documenttext-generationn<1K0 likes55 downloads3mo agoHugging Face24XReyRobert /smoke24-agentic-benchmarks Smoke24 Agentic Benchmarks Public, reproducible Terminal-Bench 2.0 Smoke24 benchmark artifacts for local RTX 3090-class agentic model serving. Why Smoke24 I created the Smoke24 subset because I needed a relatively quick benchmark that could run locally on an RTX 3090-class machine in a couple of hours. The goal is to get a practical read on model performance and stability under a real agentic Terminal-Bench workload, without paying the turnaround cost of a much… See the full description on the dataset page: https://huggingface.co/datasets/XReyRobert/smoke24-agentic-benchmarks.text-generation0 likes51 downloads1mo agoHugging Face25ArsSocratica /egora-benchmarks EgoRA Benchmark Results Comprehensive benchmark results for EgoRA (Entropy-Governed Orthogonality Regularization for Adaptation) across multiple model scales, domains, and architectures. 📦 Package: egora on PyPI 💻 Code: ArsSocratica/EgoRA on GitHub 📄 Paper: arXiv:2602.05192 🔖 DOI: 10.5281/zenodo.19398709 Dataset Structure llama-3.2-1b/, llama-3.2-3b/, llama-3.1-8b/ Fine-tuning results across 3 model scales, 2 domains (Alpaca general, Medical), 4… See the full description on the dataset page: https://huggingface.co/datasets/ArsSocratica/egora-benchmarks.text-generation1K<n<10K0 likes49 downloads6mo agoHugging Face26kevo666 /packrat-benchmarks PackRat v2 Benchmarks Version: 2.0.0 Date: 2026-04-10 Tokenizer: tiktoken cl100k_base (GPT-4 / Claude compatible) Platform: Node.js v25.6.1, Windows 11 Summary Metric Result Round-trip accuracy 100% (144/144 tests) Token savings (avg) 2.4% Token savings (best) 17.3% (path/URL-heavy files) Byte savings (avg) 2.5% Search speedup 12.03x Codebook entries 72 (auto-learned) Negative-savings entries 0 Comparison: PackRat vs MemPalace… See the full description on the dataset page: https://huggingface.co/datasets/kevo666/packrat-benchmarks.texttext-generationn<1K0 likes49 downloads6mo agoHugging Face27kylehffu /upwork-proposal-benchmarks-2026 FastBD Marketplace Proposal & Outreach Benchmarks (2026 Edition) Homepage: https://fast-bd.com Flagship Research Report: https://fast-bd.com/report-2026 Live CBR Calculator: https://fast-bd.com/cbr Official PyPI Package: fastbd-ihpi (PyPI) GitHub Repository: kylehffu/fastbd-ihpi-benchmark 1. Dataset Summary The FastBD Marketplace Proposal & Outreach Benchmarks dataset provides empirical evaluation data, unit economics benchmarks, and 52 battle-tested… See the full description on the dataset page: https://huggingface.co/datasets/kylehffu/upwork-proposal-benchmarks-2026.texttext-generationn<1K0 likes49 downloads10d agoHugging Face28Hzou9 /PCBSchemaGen-Benchmarks PCBSchemaGen Benchmarks Two benchmark suites for LLM-driven PCB schematic synthesis, from the paper PCBSchemaGen: Reward-Guided LLM Code Synthesis for Printed Circuit Board (PCB) Schematic Design with Structured Verification. Correctness in this domain is not defined by unit tests: there are no per-task golden references, and SPICE does not validate schematic-level correctness. Instead, each task is scored by a deterministic structural verifier against real-IC pin- and… See the full description on the dataset page: https://huggingface.co/datasets/Hzou9/PCBSchemaGen-Benchmarks.tabulartext-generationn<1K0 likes47 downloads3mo agoHugging Face29G3nadh /dgx-spark-benchmarks DGX Spark LLM Benchmarks First comprehensive benchmark suite for NVIDIA DGX Spark (GB10 Blackwell). Hardware GPU: NVIDIA GB10 Blackwell (1 PFLOP FP4) Memory: 128GB unified LPDDR5x (273 GB/s) CPU: 20-core ARM (10x Cortex-X925 + 10x Cortex-A725) Storage: 4TB NVMe Framework: Ollama 0.18.3 CUDA: 13.0 | Driver: 580.142 Benchmark Results Run 1 — General Inference (11 models) Model Size Prompt tok/s Gen tok/s Load Time Llama 3.1 8B 4.9 GB… See the full description on the dataset page: https://huggingface.co/datasets/G3nadh/dgx-spark-benchmarks.texttext-generationn<1K1 likes45 downloads6mo agoHugging Face30odyn-network /odyn-benchmarks Odyn Benchmarks Inference benchmark datasets and results for the Odyn Network — a distributed, OpenAI-compatible AI inference platform built on vLLM, Ray Serve, and FastAPI. Dataset Structure Prompt Profiles (data/) Four load profiles covering the full input/output token distribution space, sourced from real Odyn traffic and augmented with ShareGPT Vicuna Unfiltered: Profile Description Input tokens Output tokens Rows A Short input, Long output avg… See the full description on the dataset page: https://huggingface.co/datasets/odyn-network/odyn-benchmarks.text-generation1K<n<10K0 likes42 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.