datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
transformersVideo-MME-v2
🔥 News
2026.06.11 Videos re-encoded to H265, maintaining consistent evaluation scores. Fixed 2 incorrect MP4s & 3 mismatched URLs. Original data preserved in the original branch.
2026.05.22 Task types are now available for Q1-Q3 in coherence (logic) groups.
🤗 About This Repo
This repository contains annotation data for "Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding". It mainly consists of three… See the full description on the dataset page: https://huggingface.co/datasets/MME-Benchmarks/Video-MME-v2.daytrader-benchmarkswhat-ai-benchmarks-actually-measure
What AI Benchmarks Actually Measure: Item-Level Model Outputs and Scores for 53 Models
Item-level model responses and scores for 53 language models across the
56 benchmarks analyzed in What AI Benchmarks Actually Measure: Adapting
Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks
(Desai et al., 2026,
arxiv.org/abs/2609.08812).
We do not release the prompts from the benchmark datasets, but instead refer to them by
item ids. To regenerate the prompts from… See the full description on the dataset page: https://huggingface.co/datasets/madesai/what-ai-benchmarks-actually-measure.MiroFlow-BenchmarksThese are the benchmarking datasets used for MiroFlow Framework. More information: https://github.com/MiroMindAI/MiroThinker
speculator_benchmarksThis dataset contains dataset splits for evaluating speculative decoding algorithms on different tasks.
File:
Coding: HumanEval.jsonl
Math: math_reasoning.jsonl
Question Answering: qa.jsonl
MT_bench: question.jsonl
Retrieval-Augmented Generation: rag.jsonl
Summarization: summarization.jsonl
Translation (German to English): translation.jsonl
Writing: writing.jsonl
The data comes from two sources:
https://github.com/openai/human-eval (1). (The MIT License)… See the full description on the dataset page: https://huggingface.co/datasets/RedHatAI/speculator_benchmarks.Awesome_Spatial_VQA_BenchmarksSBI-benchmarksgraphmemix-benchmarks
GraphMemix Benchmarks
Unified multimodal memory benchmark bundles used by
GraphMemix
(arXiv:2608.26983) — four long-term
personalized memory benchmarks with their raw media assets, packaged together
for reproducible evaluation.
Benchmark
Questions
Memories
Track
Upstream license
ATM-Bench (default + hard)
1,044
11,034
memory QA over one multimodal archive
MIT
Mem-Gallery
1,711
7,944
multimodal gallery memory QA
MIT
MemEye
1,855
3,392
comics-derived memory QA… See the full description on the dataset page: https://huggingface.co/datasets/oking0197/graphmemix-benchmarks.audio_benchmarksrtx-5090-benchmarks
RTX 5090 LLM Benchmarks
Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig.
Quality Benchmarks
Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency.
Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/rtx-5090-benchmarks.benchmarks
OpenChainBench Crypto Infrastructure Benchmarks
Daily snapshots of every public benchmark on
openchainbench.com, released as
Hive-partitioned Parquet under CC-BY-4.0.
OCB measures latency, cost, coverage and accuracy of crypto
infrastructure (RPCs, oracles, bridges, data APIs, Polymarket adapters,
Hyperliquid builders). Every snapshot here mirrors the
/api/citable,
/api/stat/<slug>,
and /api/series/<slug>
JSON feeds at the time of capture.
Latest snapshot: 2026-10-06 (captured… See the full description on the dataset page: https://huggingface.co/datasets/OpenChainBench/benchmarks.vllm-inference-benchmarksSqurve-BenchmarksUI-Grounding-Benchmarks
UI-Grounding-Benchmarks
This is a collection of UI grounding benchmarks:
ScreenSpot
ScreenSpot-V2
ScreenSpot-Pro
OS-World-G
UI-Vision
Thanks for their great work!
This benchmark collection is used in the paper:
FocusUI: Efficient UI Grounding via Position-Preserving Visual Token Selection
🖼️ Project Page: https://showlab.github.io/FocusUI/
🏠 Github Repo: https://github.com/showlab/FocusUI
📝 Paper: https://arxiv.org/pdf/2601.03928
Model Zoo
Model
Backbone
🤗… See the full description on the dataset page: https://huggingface.co/datasets/yyyang/UI-Grounding-Benchmarks.Awesome_Spatial_VQA_Benchmarks_ViewSpatial-Benchexpert-rag-benchmarks
Expert RAG Benchmarks
A unified collection of four expert-level legal RAG benchmarks, exposed as six
named splits and three relational configurations: questions, documents, and
qrels.
The KCL split is named kcl_essay because Hugging Face split identifiers do not
permit hyphens; its source name remains kcl-essay.
Loading
from datasets import load_dataset
repo_id = "jinulee-v/expert-rag-benchmarks"
questions = load_dataset(repo_id, "questions", split="housing")… See the full description on the dataset page: https://huggingface.co/datasets/jinulee-v/expert-rag-benchmarks.genvsr-video-benchmarksspectre-ctrate-bimcv-500-benchmarks
v5 correction: SPECTRE embeddings re-extracted with correct resampling
The SPECTRE embeddings released in v4 (spectre_ctrate_embeddings.npz, spectre_bimcv_embeddings.npz,
their *_500_* copies, ctrate_results/, bimcv_results/ and all SPECTRE metrics derived from them) are
invalid and should not be used.
Cause: spectre-fm 0.2.1 spectre.io.resample() passes the voxel spacing to MONAI Spacing through the
deprecated affine= argument, which MONAI >= 0.9 ignores. Every scan was… See the full description on the dataset page: https://huggingface.co/datasets/chn123/spectre-ctrate-bimcv-500-benchmarks.Genomic_Benchmarks_human_enhancers_cohn
Dataset Card for "Genomic_Benchmarks_human_enhancers_cohn"
More Information needed
fastmemory-supremacy-benchmarks
FastMemory: Beyond A Million (BEAM) 10M Audit
Auditing Architectural Integrity at Scale (30 SOTA Wins)
This repository contains the official evaluation logs, simulation code, and technical whitepapers for FastMemory’s 10 Million Token BEAM Benchmark Study.
FastMemory is a sovereign, local-first memory architecture for agentic AI. Unlike traditional vector-based RAG, FastMemory utilizes Topological Isolation to achieve 100% precision in mission-critical reasoning tasks across… See the full description on the dataset page: https://huggingface.co/datasets/fastbuilderai/fastmemory-supremacy-benchmarks.D-CoT-Benchmarksbenchmarksbenchmarks
BitRouter Benchmarks
This dataset release contains the private-router Terminal-Bench 2.1 routing study and its audit evidence. July 2026 experiments are preserved under v0/; the current study, raw Harbor traces, run-scoped BitRouter SQLite files, sanitized production-usage backfill, reproducible joins, and Parquet tables are under terminal-bench-2.1/router-random-study/.
Main result
All valid-case metrics below use the same strict intersection of 80 tasks. “Harbor… See the full description on the dataset page: https://huggingface.co/datasets/BitRouterAI/benchmarks.lra-benchmarksdata-preprocessing-automl-benchmarks
Data Preprocessing AutoML Benchmarks
This repository contains text classification datasets with known data quality issues for preprocessing research in AutoML.
Usage
Load a specific dataset configuration like this:
from datasets import load_dataset
# Example for loading the TREC dataset
dataset = load_dataset("MothMalone/data-preprocessing-automl-benchmarks", "trec")
Available Datasets
Below are the details for each dataset configuration available in this… See the full description on the dataset page: https://huggingface.co/datasets/MothMalone/data-preprocessing-automl-benchmarks.Genomic_Benchmarks_human_nontata_promoters
Dataset Card for "Genomic_Benchmarks_human_nontata_promoters"
More Information needed
LEM-benchmarks
LEM-benchmarks
Canonical 8-PAC benchmark results for the Lemma model family.
This dataset is an aggregated store of per-round evaluation data produced by
lthn/LEM-Eval. Every row
represents one model's answer to one question in one round of a paired A/B
run against its unmodified base, and the dataset grows monotonically as more
workers contribute — different machines, different sampling states, different
hardware paths — which is the whole point of 8-PAC: multiple independent… See the full description on the dataset page: https://huggingface.co/datasets/lthn/LEM-benchmarks.Genomic_Benchmarks_demo_coding_vs_intergenomic_seqs
Dataset Card for "Genomic_Benchmarks_demo_coding_vs_intergenomic_seqs"
More Information needed
mg3d-ctrate-bimcv-500-benchmarks
MG-3D Swin-B (MIA 2026) Benchmark on CT-RATE 500 & BIMCV-R 500 Cohorts (v4 Definitive Release)
This repository contains the definitive out-of-fold benchmark results, extracted 3D embeddings, probability predictions, affine/HU calibration manifests, and reproducibility code for MG-3D Swin-B (MIA 2026) evaluated across two independent clinical 3D Computed Tomography cohorts:
CT-RATE Validation Cohort (500 series, 428 unique patients, 463 unique studies)
BIMCV-R Cohort (317… See the full description on the dataset page: https://huggingface.co/datasets/chn123/mg3d-ctrate-bimcv-500-benchmarks.
