datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
explicit-edit-benchmark
Explicit Edit Benchmark
226 deterministic exact-edit tasks, run by different agents, harnesses, models and configurations. Every observation records what the harness did and whether the resulting files matched byte for byte.
Source code and benchmark runner: GitHub — Explicit Edit Benchmark
Open the interactive Explorer to compare agents, harnesses, models, versions, reasoning modes, correctness, recovery, time, cost and tokens.
Leaderboard by model route
Score v2… See the full description on the dataset page: https://huggingface.co/datasets/alexshpunt/explicit-edit-benchmark.daytrader-benchmarkswhat-ai-benchmarks-actually-measure
What AI Benchmarks Actually Measure: Item-Level Model Outputs and Scores for 53 Models
Item-level model responses and scores for 53 language models across the
56 benchmarks analyzed in What AI Benchmarks Actually Measure: Adapting
Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks
(Desai et al., 2026,
arxiv.org/abs/2609.08812).
We do not release the prompts from the benchmark datasets, but instead refer to them by
item ids. To regenerate the prompts from… See the full description on the dataset page: https://huggingface.co/datasets/madesai/what-ai-benchmarks-actually-measure.regx-benchmark
RegX
Cross-Domain Multi-View Point Cloud Registration Benchmark
RegX evaluates multi-view point cloud registration across scales spanning nine orders
of magnitude — nanometre-scale microscopy to kilometre-scale airborne maps — and sensors
never designed to be compared: clinical colonoscopes, RGB-D cameras, spinning and
solid-state LiDAR, terrestrial and airborne laser scanners.
Most registration benchmarks fix one sensor and one scale. RegX asks a narrower question
instead: does… See the full description on the dataset page: https://huggingface.co/datasets/YuePanEdward/regx-benchmark.results_public
Dataset Card for "resultspublic"
More Information needed
STRABLE-benchmark
STRABLE: Benchmarking Tabular Machine Learning with Strings
This dataset card describes the STRABLE benchmark, a comprehensive suite designed for evaluating machine learning models on tabular data containing strings.
Dataset Description
Benchmarking tabular data has revealed the benefit of dedicated architectures, pushing the state of the art. However, real-world tables often contain string entries beyond pure numbers, a setting that has been understudied due to a… See the full description on the dataset page: https://huggingface.co/datasets/inria-soda/STRABLE-benchmark.multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR
PreFLMR M2KR Dataset Card
Dataset details
Dataset type:
M2KR is a benchmark dataset for multimodal knowledge retrieval. It contains a collection of tasks and datasets for training and evaluating multimodal knowledge retrieval models.
We pre-process the datasets into a uniform format and write several task-specific prompting instructions for each dataset. The details of the instruction can be found in the paper. The M2KR benchmark contains three types of tasks:… See the full description on the dataset page: https://huggingface.co/datasets/BByrneLab/multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR.funes-handoff-recall-benchmark
handover-vs-recall
A long investigation bloats an agent session until each new turn costs more to carry the context than to
do the work. Switching to a fresh session avoids that — but the findings have to travel somehow, and the
ways of moving them differ in cost. This benchmark measures those ways, as cost per successful task,
on tasks that genuinely require the prior investigation:
arm
channel
A branch-only
switch, carry nothing — the fresh session re-derives the… See the full description on the dataset page: https://huggingface.co/datasets/dacorvo/funes-handoff-recall-benchmark.benchmark-llms-landuse-relevance
Land-use relevance benchmark
v3-multilingual · 85 languages x 300 items/language ·
25,500 items · binary yes/no labels.
Code
Package version recorded in run metadata: 0.2.0 (some runs lack version metadata).
Task and prompt
Does a sentence describe a place's land or environment in ways visible to satellites?
English prompt · greedy decoding · seed 0 · max_new_tokens=4096 ·
bfloat16 · batch varies by model.
unsloth/Qwen3.8-27B-GGUF@UD-IQ2_XXS runs the UD-IQ2_XXS… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/benchmark-llms-landuse-relevance.geoparser-benchmark-results
Geoparser benchmark results
Geoparsing pipelines from geoparser scored on English and
multilingual corpora, run on Grid'5000 (one Tesla T4).
Benchmarks
Benchmark
Languages
Docs
Toponyms
Source
GeoVirus
en
229
2167
WikiNews articles on epidemics (Gritta et al., 2018).
HIPE-2020
de, en, fr
129
1516
Historical Swiss, Luxembourgish and American newspapers, OCR.
NewsEye
de, fi, fr, sv
77
1772
Historical European newspapers, OCR (HIPE-2022).
TopRes19th… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/geoparser-benchmark-results.ai-humanizer-benchmark
AI Humanizer Benchmark: AI humanizers tested against 7 AI detectors (October 2026)
AI Humanizer Benchmark measures how well AI humanizers rewrite AI-generated text so that AI detectors classify it as human-written, and how much meaning and readability the rewrite loses. In each monthly cycle, 11 AI humanizers rewrite the same 33 source texts across 7 writing categories, using each tool's default settings. Each output gets three kinds of score: 7 AI detectors (GPTZero… See the full description on the dataset page: https://huggingface.co/datasets/ai-humanizer-benchmark/ai-humanizer-benchmark.stan-benchmarkinvoice-extraction-benchmark
Invoice Extraction Benchmark v1
A synthetic test set for invoice data extraction (invoice OCR, intelligent document
processing, accounts-payable capture): 181 documents with answer keys and a scorer.
Run any invoice reader over the documents, write its output as one JSON file, and score it
field by field.
Source, scorer and generator: https://github.com/DrewKraken/invoice-extraction-benchmark
(this dataset is a mirror of corpus/v1/ there; the GitHub repository is canonical).
Who… See the full description on the dataset page: https://huggingface.co/datasets/drew-ipp/invoice-extraction-benchmark.tabular-benchmark
Tabular Benchmark
Dataset Description
This dataset is a curation of various datasets from openML and is curated to benchmark performance of various machine learning algorithms.
Repository: https://github.com/LeoGrin/tabular-benchmark/community
Paper: https://hal.archives-ouvertes.fr/hal-03723551v2/document
Dataset Summary
Benchmark made of curation of various tabular data learning tasks, including:
Regression from Numerical and Categorical Features… See the full description on the dataset page: https://huggingface.co/datasets/inria-soda/tabular-benchmark.mtg-urna-benchmark
38,627 Magic: The Gathering cards, one per oracle id, the scan and the rules text of each, packed into single .urna files that answer text and image queries from memory-mapped bytes: no server, no Python at read time. Ten such files live here. They carry the same cards, the same text and the same content hash; what differs is how the 4 GB of JPEG was encoded inside, and which image models embedded it.
The file to start with is release/v0.3/stills-5models. It is the only one with the models… See the full description on the dataset page: https://huggingface.co/datasets/brennercruvinel/mtg-urna-benchmark.multilingual-tts-benchmark
Multilingual Speech Benchmark for Zero-Shot TTS
A voice-cloning and intelligibility benchmark for 8 language subsets,
with 10,100 examples selected from Common Voice 17.0. Each example supplies
a speaker reference and an independently selected target text, with human
recordings as WER/CER and speaker-similarity anchors when the corresponding
audio is available. Corpus WER and the WavLM-FT evaluation follow
seed-tts-eval.
Version 3.1. Adds ru, kk through the same S1–S4 selection… See the full description on the dataset page: https://huggingface.co/datasets/nineninesix/multilingual-tts-benchmark.lighting-invariant-bedroom-perception-robustness-benchmark
Lighting-Invariant Bedroom Perception & Robustness Benchmark
Generated by datapack-import.ts
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is… See the full description on the dataset page: https://huggingface.co/datasets/physicl/lighting-invariant-bedroom-perception-robustness-benchmark.LEXam
LEXam: Benchmarking Legal Reasoning on 340 Law Exams
A diverse, rigorous evaluation suite for legal AI from Swiss, EU, and international law examinations.
Paper | Website & Leaderboard | GitHub Repository
🔥 News
[2026/01] Our paper has been accepted to ICLR 2026!
[2025/12] We reorganized all multiple-choice questions into four separate files, mcq_4_choices (n = 1,655), mcq_8_choices (n = 1,463), mcq_16_choices (n = 1,028), and mcq_32_choices (n = 550), all… See the full description on the dataset page: https://huggingface.co/datasets/LEXam-Benchmark/LEXam.video-benchmark-resultsxplanner-benchmark
XPlanner-Benchmark
XPlanner-Benchmark is the portable release of the X-Planner 1,500-episode evaluation benchmark. It contains synchronized multi-view robot-manipulation videos and the episode-level task, subtask, action, scene, duration, and complexity metadata used by X-Planner.
Contents
1,500 episodes
3,490 MP4 video references
167 source dataset identifiers
525 unique task names
31 task classes
41 inferred atomic action labels
22.70 total hours of episode… See the full description on the dataset page: https://huggingface.co/datasets/x-square-robot/xplanner-benchmark.rtx-5090-benchmarks
RTX 5090 LLM Benchmarks
Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig.
Quality Benchmarks
Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency.
Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/rtx-5090-benchmarks.Krea-2-Turbo-Checkpoint-Format-Benchmark
Krea 2 Turbo ComfyUI Format Fidelity Benchmark
This release is a paired, deterministic comparison of eight Krea 2 Turbo checkpoint formats in ComfyUI: BF16, FP8 Scaled, INT8 ConvRot, MXFP8, NVFP4, INT4 ConvRot W4A4, GGUF Q8_0, and GGUF Q4_K_M. It contains 240 scored 1024×1024 images, saved float32 decoded tensors and final latents, every denoising trajectory, raw metric tables, telemetry, statistical comparisons, and reproduction code.
Main result
BF16 is the… See the full description on the dataset page: https://huggingface.co/datasets/Merserk/Krea-2-Turbo-Checkpoint-Format-Benchmark.b3-agent-security-benchmark-weak[paper] [blogpost] [game]
b3 AI Security Benchmark: Breaking Agent Backbones
Highly contextalized prompt injections crowd-sourced during the Agent Breaker Challenge.
This is a low-quality version of the data behind Breaking Agent Backbones: Evaluating the Security
of Backbone LLMs in AI Agents.
The high quality dataset was used to evaluate the security of more than 30 LLMs.
Dataset Summary
Purpose: This dataset contains crowdsourced adversarial attacks… See the full description on the dataset page: https://huggingface.co/datasets/Lakera/b3-agent-security-benchmark-weak.jetson-non-reasoning-benchmark-ollama-15w
Tiny LLM Benchmark — Jetson Orin Nano Super 8GB
Date: 2026-06-07 02:35Backends: ollamaSweep: prompt ∈ {128,512,1024,2048} tok × gen ∈ {64,128,256} tokArtifacts: /home/yuvrajsingh/Desktop/benchmark/smolbenchmark/non-reasoning-models/artifacts/blog-all-20260606-0139-15w
Full Results — ollama
Power = VDD_CPU_GPU_CV avg over aiperf window.
Model
Quant
ISL
OSL
OSL mis%
TTFT avg
p50
p90
p99
T2T avg
p50
p90
p99
ITL avg
p50
p90
p99
Tok/s
Req/s
E2E avg
p50
p90
p99… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/jetson-non-reasoning-benchmark-ollama-15w.oparq-benchmarks
oparq benchmark results
source code · PyPI package · benchmark inputs
Results only: 21 datasets, 967,938,981 input rows,
986 source files. No original source rows are distributed.
The viewer rows describe algorithm measurements, not individual source events.
Full-corpus results sort all input rows globally using fixed writer settings;
only key planning is sampled. The separate Arrow/DuckDB comparison sorts
within each physical file using saved keys, and must not be conflated… See the full description on the dataset page: https://huggingface.co/datasets/Poupou/oparq-benchmarks.jetson-non-reasoning-benchmark-ollama-7w
Tiny LLM Benchmark — Jetson Orin Nano Super 8GB
Date: 2026-06-09 02:38Backends: ollamaSweep: prompt ∈ {128,512,1024,2048} tok × gen ∈ {64,128,256} tokArtifacts: /home/yuvrajsingh/Desktop/benchmark/smolbenchmark/non-reasoning-models/artifacts/blog-all-20260607-0403-7w
Full Results — ollama
Power = VDD_CPU_GPU_CV avg over aiperf window.
Model
Quant
ISL
OSL
OSL mis%
TTFT avg
p50
p90
p99
T2T avg
p50
p90
p99
ITL avg
p50
p90
p99
Tok/s
Req/s
E2E avg
p50
p90
p99… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/jetson-non-reasoning-benchmark-ollama-7w.benchmarks
OpenChainBench Crypto Infrastructure Benchmarks
Daily snapshots of every public benchmark on
openchainbench.com, released as
Hive-partitioned Parquet under CC-BY-4.0.
OCB measures latency, cost, coverage and accuracy of crypto
infrastructure (RPCs, oracles, bridges, data APIs, Polymarket adapters,
Hyperliquid builders). Every snapshot here mirrors the
/api/citable,
/api/stat/<slug>,
and /api/series/<slug>
JSON feeds at the time of capture.
Latest snapshot: 2026-10-10 (captured… See the full description on the dataset page: https://huggingface.co/datasets/OpenChainBench/benchmarks.bonsai-jetson-benchmark-15w
Bonsai Jetson Benchmark — 15W
Platform: NVIDIA Jetson Orin Nano Super 8GB · Power mode: 15W
Backend: llama.cpp build-jetson · CUDA · -ngl 99
Sweep: prompt ∈ {256, 512, 1024, 2048} tok × gen ∈ {128, 256, 512} tok · 20 reqs/combo
Status: Complete — 57 combos (5 models × 12 prompt/gen configs)
Key metric: tok/J = output tok/s ÷ VDD_CPU_GPU_CV (W)
Models
Model
Quant
Size
Bonsai-1.7B
Q1_0 (1-bit)
~237 MB
Bonsai-4B
Q1_0 (1-bit)
~540 MB
Bonsai-8B
Q1_0… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/bonsai-jetson-benchmark-15w.jetson-non-reasoning-benchmark-ollama-25w
Tiny LLM Benchmark — Jetson Orin Nano Super 8GB
Date: 2026-06-23 06:04Backends: ollamaSweep: prompt ∈ {128,512,1024,2048} tok × gen ∈ {64,128,256} tokArtifacts: /home/yuvrajsingh/Desktop/benchmark/smolbenchmark/benchmark-jetson-nano-orin-super/non-reasoning-models/artifacts/blog-all-20260622-0159-25w
Full Results — ollama
Power = VDD_CPU_GPU_CV avg over aiperf window.
Model
Quant
ISL
OSL
OSL mis%
TTFT avg
p50
p90
p99
T2T avg
p50
p90
p99
ITL avg
p50
p90
p99… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/jetson-non-reasoning-benchmark-ollama-25w.jetson-non-reasoning-benchmark-7w
Tiny LLM Benchmark — Jetson Orin Nano Super 8GB (7W / nvpmodel -m 3)
Date: 2026-05-28 18:03Backend: llama.cpp CUDA (-ngl 99) Concurrency: 1Sweep: prompt in {128,512,1024,2048} tok × gen in {64,128,256} tok
Artifacts: https://huggingface.co/datasets/YuvrajSingh9886/jetson-non-reasoning-benchmark-7w
Note: tok/J computed from per-run start_time/end_time in each aiperf JSON
Full Results
Power = VDD_CPU_GPU_CV average over each aiperf run window (per-run timestamps… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/jetson-non-reasoning-benchmark-7w.
