datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
explicit-edit-benchmark
Explicit Edit Benchmark
226 deterministic exact-edit tasks, run by different agents, harnesses, models and configurations. Every observation records what the harness did and whether the resulting files matched byte for byte.
Source code and benchmark runner: GitHub — Explicit Edit Benchmark
Open the interactive Explorer to compare agents, harnesses, models, versions, reasoning modes, correctness, recovery, time, cost and tokens.
Leaderboard by model route
Score v2… See the full description on the dataset page: https://huggingface.co/datasets/alexshpunt/explicit-edit-benchmark.LEXam
LEXam: Benchmarking Legal Reasoning on 340 Law Exams
A diverse, rigorous evaluation suite for legal AI from Swiss, EU, and international law examinations.
Paper | Website & Leaderboard | GitHub Repository
🔥 News
[2026/01] Our paper has been accepted to ICLR 2026!
[2025/12] We reorganized all multiple-choice questions into four separate files, mcq_4_choices (n = 1,655), mcq_8_choices (n = 1,463), mcq_16_choices (n = 1,028), and mcq_32_choices (n = 550), all… See the full description on the dataset page: https://huggingface.co/datasets/LEXam-Benchmark/LEXam.rtx-5090-benchmarks
RTX 5090 LLM Benchmarks
Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig.
Quality Benchmarks
Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency.
Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/rtx-5090-benchmarks.b3-agent-security-benchmark-weak[paper] [blogpost] [game]
b3 AI Security Benchmark: Breaking Agent Backbones
Highly contextalized prompt injections crowd-sourced during the Agent Breaker Challenge.
This is a low-quality version of the data behind Breaking Agent Backbones: Evaluating the Security
of Backbone LLMs in AI Agents.
The high quality dataset was used to evaluate the security of more than 30 LLMs.
Dataset Summary
Purpose: This dataset contains crowdsourced adversarial attacks… See the full description on the dataset page: https://huggingface.co/datasets/Lakera/b3-agent-security-benchmark-weak.bonsai-jetson-benchmark-15w
Bonsai Jetson Benchmark — 15W
Platform: NVIDIA Jetson Orin Nano Super 8GB · Power mode: 15W
Backend: llama.cpp build-jetson · CUDA · -ngl 99
Sweep: prompt ∈ {256, 512, 1024, 2048} tok × gen ∈ {128, 256, 512} tok · 20 reqs/combo
Status: Complete — 57 combos (5 models × 12 prompt/gen configs)
Key metric: tok/J = output tok/s ÷ VDD_CPU_GPU_CV (W)
Models
Model
Quant
Size
Bonsai-1.7B
Q1_0 (1-bit)
~237 MB
Bonsai-4B
Q1_0 (1-bit)
~540 MB
Bonsai-8B
Q1_0… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/bonsai-jetson-benchmark-15w.benchmark-radar
Benchmark Radar Dataset
Overview
Benchmark Radar is a living registry, search engine, and discovery pipeline
for AI evaluation benchmarks. This dataset mirrors the full-corpus findings of the
Benchmark Radar technical report ("Benchmark Radar: Daily Discovery and
Full-Corpus Search Across the AI Evaluation Landscape", arXiv:2609.11115).
As described in the paper's Two Input Paths framework, Benchmark Radar
combines two complementary systems:… See the full description on the dataset page: https://huggingface.co/datasets/ktwu01/benchmark-radar.os-omni-benchmark
OS-Omni Benchmark
OS-Omni is a cross-platform benchmark for evaluating agents that operate graphical operating-system environments. This dataset repository contains the static benchmark task definitions and supporting assets used to configure and evaluate OS-Omni tasks.
Contents
data/tasks.parquet: tabular task index for Hugging Face Dataset Viewer and Croissant generation.
data/tasks.jsonl: JSON Lines copy of the same task index.
metadata/tasks.parquet: duplicate task… See the full description on the dataset page: https://huggingface.co/datasets/OSOmni/os-omni-benchmark.local-llm-benchmark
Local LLM Benchmark — Technical and Uncensored Behavior (NVIDIA RTX 5070 Ti 16GB)
English | 简体中文 | 繁體中文 | 한국어 | Español | 日本語 | हिन्दी | Русский | Português | తెలుగు | Français | Deutsch | Italiano | Tiếng Việt | العربية | اردو | বাংলা | فارسی | Română | Türkçe
Manual evaluation results of local GGUF model variants on a single consumer machine,
combining two fully independent benchmarks:
technical/
uncensored/
Measures
capability: coding, systems, networking, DB, agents… See the full description on the dataset page: https://huggingface.co/datasets/nanimani/local-llm-benchmark.legal-llm-benchmark
Legal LLM Benchmark Dataset
Safety-Utility Trade-offs in Legal AI: An LLM Evaluation Across 12 Models
Quick Start
from datasets import load_dataset
# Core datasets
questions = load_dataset("marvintong/legal-llm-benchmark", "questions")
phase1_evals = load_dataset("marvintong/legal-llm-benchmark", "phase1_evaluations")
phase3_evals = load_dataset("marvintong/legal-llm-benchmark", "phase3_evaluations")
# Additional helpful datasets
contracts =… See the full description on the dataset page: https://huggingface.co/datasets/marvintong/legal-llm-benchmark.story_writing_benchmark
Story Evaluation Dataset
This dataset contains stories generated by Large Language Models (LLMs) across multiple languages, with comprehensive quality evaluations. It was created to train and benchmark models specifically on creative writing tasks.
This benchmark evaluates an LLM's ability to generate high-quality short stories based on simple prompts like "write a story about X with n words." It is similar to TinyStories but targets longer-form and more complex content, focusing… See the full description on the dataset page: https://huggingface.co/datasets/lars1234/story_writing_benchmark.Cognitive_Atrophy_Benchmark
Cognitive Atrophy Benchmark — LLM Responses Across Four Mental-Health Conversation Datasets
This dataset releases the LLM-response component of the Cognitive Atrophy Benchmark: five large language models prompted under identical conditions across four mental-health conversation datasets. It is a building block for a forthcoming evaluation framework that quantifies cognitive atrophy — the gradual erosion of users' own reasoning, recall, and decisional autonomy when an LLM… See the full description on the dataset page: https://huggingface.co/datasets/abadawi/Cognitive_Atrophy_Benchmark.rtx-5090-benchmarks
RTX 5090 LLM Benchmarks
Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig.
Quality Benchmarks
Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency.
Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/omegaprime669/rtx-5090-benchmarks.auditable-fincrime-benchmark
Auditable Financial-Crime Compliance Benchmark
A matched-twin benchmark of auditable explanations for financial-crime compliance, spanning two domains — anti-money-laundering (AML) and market abuse. Every scenario is judged not by a yes/no label but against a four-element rubric: the mechanism of the abuse, the rule that is defeated (loophole in AML, breach in market abuse), the distinguishing facts that separate a violation from its innocent twin, and the decision (REPORT /… See the full description on the dataset page: https://huggingface.co/datasets/Anonymousbabygoat/auditable-fincrime-benchmark.SciCode-Runnable-Benchmark-Reviewedqwen36-cross-platform-benchmark
Qwen3.6-35B-A3B cross-platform benchmark dataset
Complete sanitized evidence for the Qwen3.6 GX10 versus M2 Max benchmark.
Headline results
At 128K, median cold-prompt TTFT was 45.70 s on GX10 and 552.79 s on M2 Max, a 12.10x difference.
Near 256K, GX10 completed and strictly passed 12/12 requests. M2 Max completed 9/12 and strictly passed 7/9 completed requests.
On the same Q4_K_M coding control, MTP improved median decode 26.87% on GX10 CUDA and regressed… See the full description on the dataset page: https://huggingface.co/datasets/tekosML/qwen36-cross-platform-benchmark.multilingual_tokenizer_benchmark
Multilingual Tokenizer Benchmark
More details of each subset like word count, character count, original sources, etc, can be found in the dataset_meta.yaml file in the repository root.
Natural language word count functions
Download spacy models
pip install ntlk spacy pygments underthesea camel-tools
python -m spacy download ko_core_news_sm
python -m spacy download ja_core_news_sm
python -m spacy download zh_core_web_sm
import nltk
nltk.download('punkt_tab')… See the full description on the dataset page: https://huggingface.co/datasets/eduagarcia/multilingual_tokenizer_benchmark.vault-benchmark-v1
APort Vault Benchmark v1
4,371 attacks written by humans in 1,128 sessions of a public competition
(March to August 2026), replayed against 14 language models from 8 labs acting
as a simulated bank teller (no real money moves), in two conditions: model
alone, and behind a pre-action authorization layer.
Papers
APort Vault: Benchmarking AI Agent Payment Authorization with the Open Agent Passport (Paper 2). The replay benchmark and results reported in this dataset.… See the full description on the dataset page: https://huggingface.co/datasets/aporthq/vault-benchmark-v1.indian-responsible-ai-benchmark
Indian Responsible AI Benchmark
A comprehensive benchmark for evaluating responsible AI behavior in Indian contexts — covering 212 adversarial and safety-critical prompts across 22 evaluation categories, 10 Indian language regions, and 8 Responsible AI dimensions.
Why This Benchmark?
Most AI safety benchmarks are US/Western-centric. Indian users face unique challenges:
Caste dynamics not captured by Western bias benchmarks
India/US context confusion (models… See the full description on the dataset page: https://huggingface.co/datasets/responsible-ai-labs/indian-responsible-ai-benchmark.NPM-Artifact-Explanation-Benchmark
NPM-Artifact-Explanation-Benchmark
English
NPM-Artifact-Explanation-Benchmark is a cross-category multimodal corpus and benchmark resource for Chinese cultural artifact understanding and explanation.
This release contains 28,826 cleaned artifact records derived from National Palace Museum source records' opendata (https://digitalarchive.npm.gov.tw/opendata/). Each record includes structured artifact metadata, image URLs, source record URLs, and human-written… See the full description on the dataset page: https://huggingface.co/datasets/shunanhe/NPM-Artifact-Explanation-Benchmark.PDB-Single
PDB-Single: Precise Debugging Benchmarking — single-line bug set
📄 Paper ·
💻 Code ·
🌐 Project page ·
🏆 Leaderboard
PDB-Single is the single-line bug set of the PDB (Precise Debugging Benchmarking) evaluation suite. Every example pairs a ground-truth program with a synthesized buggy version plus a line-level edit script (gt_diff) that encodes the minimal correct fix.
Source datasets: BigCodeBench + LiveCodeBench
Sibling datasets:… See the full description on the dataset page: https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Single.emotion-negotiation-benchmarks
Emotion-Aware LLM Negotiation Benchmarks
Four high-stakes, edge-deployable negotiation benchmarks — the official evaluation suite for our research program on emotion-aware LLM agents. Each benchmark targets a distinct domain where (a) LLM-vs-LLM negotiation has real-world consequences, and (b) on-device deployment of small language models matters for privacy and latency.
The benchmarks were originally introduced with EmoMAS (ACL 2026 Main, top 9% of 12,148 submissions) and are… See the full description on the dataset page: https://huggingface.co/datasets/humanlong/emotion-negotiation-benchmarks.singapore-legal-ai-benchmark
Singapore Legal AI Benchmark
Public research release of 102 Singapore legal research questions, model
responses from 6 systems, and overlapping grades on five dimensions.
Headline metrics are overlapping binary flags, not a ranking and not a
partition of 100%.
Interactive explorer
Open the explorer →
— comparison table, category heatmap, per-question comparison, and every answer
with its sources and grades.
(Space page)
Overall (n = 612)… See the full description on the dataset page: https://huggingface.co/datasets/JonathanSu/singapore-legal-ai-benchmark.l4-gpu-llm-benchmark-leaderboard
🚀 Local LLM Serving & Quality Benchmark Leaderboard (NVIDIA L4 24GB)
An exhaustive, reproducible benchmark study measuring real-world serving performance (TTFT, TPOT, throughput, peak VRAM, energy consumption, and cost) alongside rigorous task quality gates (HumanEval+, MMLU-Pro, BFCL v4 tool calling, and RULER needle retrieval) for open-weight LLMs on a single NVIDIA L4 24GB GPU.
📊 Executive Summary & Key Takeaways
⚡ Best Throughput & Coding Workhorse:… See the full description on the dataset page: https://huggingface.co/datasets/mayank-dubey-ai/l4-gpu-llm-benchmark-leaderboard.math_benchmark_test_saturation
LLM Leaderboard Data for Hendrycks MATH Dataset (2022–2024)
This dataset aggregates yearly performance (2022–2024) of large language models (LLMs) on the Hendrycks MATH benchmark. It is specifically compiled to explore performance evolution, benchmark saturation, parameter scaling trends, and evaluation metrics of foundation models solving complex math word problems.
Original source data: Math Word Problem Solving on MATH (Papers with Code)
About Hendrycks' MATH… See the full description on the dataset page: https://huggingface.co/datasets/nlile/math_benchmark_test_saturation.hallucination-autopsy-benchmark
Hallucination Autopsy Benchmark
A unified, standardized benchmark for cross-model, cross-parameter analysis of LLM hallucination phenomena.
Overview
This dataset merges multiple hallucination detection benchmarks into a single standardized schema, enabling systematic etiological analysis of why and how different LLM architectures hallucinate under specific configurations.
Version: 3.0.0Total Records: 69,002Base Records: 69,002Augmented Records: 0Hallucination… See the full description on the dataset page: https://huggingface.co/datasets/OiQ/hallucination-autopsy-benchmark.compile-benchmark
CompilingThings Compile Benchmark for MQL5®
This release evaluates compile success of generated MQL5 on two private held-out sets. The first is 300 Expert Advisor prompts, run on four arms: the base model, two tuned local models and one frontier API model. The second is 200 non-EA prompts (include files, custom indicators, scripts and services, 50 each), run on the three local arms, plus a stability re-run of one of them. The holdout results are attested, not fully verifiable:… See the full description on the dataset page: https://huggingface.co/datasets/CompilingThings/compile-benchmark.hemmingway-1-omlx-quantization-benchmark-v1
Hemmingway-1 oMLX Quantization Benchmark
This is the public-safe benchmark package for the Hemmingway-1 oMLX
quantization study on Apple Silicon.
Altworld developed and published
Hemmingway-1. Bobby Pierce
published these quantizations and the evaluation package. The
collection
links the upstream model and all six builds.
Analysis revision 2, corrected on 2026-09-22, fixes A/B attribution and matching
across reversed packets. Read CORRECTION.md before using the
aggregate… See the full description on the dataset page: https://huggingface.co/datasets/sixstringzen/hemmingway-1-omlx-quantization-benchmark-v1.qwen3.8-27b-medical-benchmarks
Qwen3.8-27B — Medical Benchmark Evaluation
Evaluation report for Qwen/Qwen3.8-27B across 32 medical benchmark configurations
(27,312 scored items), under a single fixed inference configuration, with two grading regimes:
deterministic and model-graded.
The model was evaluated as released. No weights were modified.
Field
Value
Model under evaluation
Qwen/Qwen3.8-27B
Evaluation dates
2026-09-27 to 2026-09-30 (UTC)
Benchmark configurations
32 (24 deterministic, 8… See the full description on the dataset page: https://huggingface.co/datasets/himajahealth/qwen3.8-27b-medical-benchmarks.llmfit-benchmarks
llmfit Real-World LLM Inference Benchmarks
An evolving dataset of real-world LLM inference measurements across consumer, workstation, datacenter, and unified-memory hardware. It combines benchmarks from an external community source with community benchmarks contributed directly to llmfit.
The initial release contains 1,501 normalized observations:
1,010 unique external-community observations from the repository's 2026-08-10 snapshot.
491 llmfit-community observations from 61… See the full description on the dataset page: https://huggingface.co/datasets/axjns/llmfit-benchmarks.opencode-ai-benchmark
🌐 OpenCode / Antigravity Protocol (October 2026)
Professional AI Engineering & Cybersecurity Benchmark
Deep Reasoning • Human-Like Engineering Judgment • Adversarial Traps • Zero Fabrication
Live Interactive Leaderboard • Executive Report • 120 Skills Taxonomy • Evaluation Protocol • Quickstart
[!WARNING]
⚠️ EXPERIMENTAL TRIAL RELEASE (VERSI UJI COBA)
Research Preview Notice: This benchmark dataset, leaderboard, and… See the full description on the dataset page: https://huggingface.co/datasets/Kicaulah/opencode-ai-benchmark.
