datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
frames-benchmark
FRAMES: Factuality, Retrieval, And reasoning MEasurement Set
FRAMES is a comprehensive evaluation dataset designed to test the capabilities of Retrieval-Augmented Generation (RAG) systems across factuality, retrieval accuracy, and reasoning.
Our paper with details and experiments is available on arXiv: https://arxiv.org/abs/2409.12941.
Dataset Overview
824 challenging multi-hop questions requiring information from 2-15 Wikipedia articles
Questions span diverse topics… See the full description on the dataset page: https://huggingface.co/datasets/google/frames-benchmark.benchmark-llms-landuse-relevance
Land-use relevance benchmark
v3-multilingual · 85 languages x 300 items/language ·
25,500 items · binary yes/no labels.
Code
Package version recorded in run metadata: 0.2.0 (some runs lack version metadata).
Task and prompt
Does a sentence describe a place's land or environment in ways visible to satellites?
English prompt · greedy decoding · seed 0 · max_new_tokens=4096 ·
bfloat16 · batch varies by model.
unsloth/Qwen3.8-27B-GGUF@UD-IQ2_XXS runs the UD-IQ2_XXS… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/benchmark-llms-landuse-relevance.geoparser-benchmark-results
Geoparser benchmark results
Geoparsing pipelines from geoparser scored on English and
multilingual corpora, run on Grid'5000 (one Tesla T4).
Benchmarks
Benchmark
Languages
Docs
Toponyms
Source
GeoVirus
en
229
2167
WikiNews articles on epidemics (Gritta et al., 2018).
HIPE-2020
de, en, fr
129
1516
Historical Swiss, Luxembourgish and American newspapers, OCR.
NewsEye
de, fi, fr, sv
77
1772
Historical European newspapers, OCR (HIPE-2022).
TopRes19th… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/geoparser-benchmark-results.L-FAME
L-FAME: Longitudinal Focused Attention Meditation EEG Dataset and Benchmark
A longitudinal 64-channel EEG dataset and benchmark for studying focused attention meditation (FAM) and how its neural signatures evolve across a six-week training period. 74 healthy adults were recorded at a pre-intervention baseline; 44 of them returned for a post-intervention follow-up. Three FAM techniques are systematically compared: Hare Krishna mantra (HK), SA-TA-NA-MA mantra (SA), and Breath… See the full description on the dataset page: https://huggingface.co/datasets/L-FAME-Dataset-Benchmark/L-FAME.tabular-benchmark
Tabular Benchmark
Dataset Description
This dataset is a curation of various datasets from openML and is curated to benchmark performance of various machine learning algorithms.
Repository: https://github.com/LeoGrin/tabular-benchmark/community
Paper: https://hal.archives-ouvertes.fr/hal-03723551v2/document
Dataset Summary
Benchmark made of curation of various tabular data learning tasks, including:
Regression from Numerical and Categorical Features… See the full description on the dataset page: https://huggingface.co/datasets/inria-soda/tabular-benchmark.finance_agent_benchmark
Finance Agent Benchmark Dataset
We present the Finance Agent Benchmark, featuring challenging and diverse real-world finance research problems which require LLMs to perform complex analysis with the use of of recent SEC filings.
We construct the benchmark using a taxonomy of nine financial task categories, developed in consultation with experts from banks, hedge funds, and private equity firms. The dataset includes 537 expert-authored questions, covering tasks from information… See the full description on the dataset page: https://huggingface.co/datasets/vals-ai/finance_agent_benchmark.rtx-5090-benchmarks
RTX 5090 LLM Benchmarks
Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig.
Quality Benchmarks
Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency.
Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/rtx-5090-benchmarks.vllm-inference-benchmarksAntibody_Binding_Benchmark_DatasetWe introduce AbBiBench (Antibody Binding Benchmarking), a benchmarking framework for optimizing antibody binding affinity. This dataset contains the sequences of mutants, experimentally measured affinity values, and the structures of antigen-antibody complexes.
Mutant structure files for AbBiBench dataset is available at https://zenodo.org/records/16557372
b3-agent-security-benchmark-weak[paper] [blogpost] [game]
b3 AI Security Benchmark: Breaking Agent Backbones
Highly contextalized prompt injections crowd-sourced during the Gandalf Agent Breaker Challenge.
This is a low-quality version of the data behind Breaking Agent Backbones: Evaluating the Security
of Backbone LLMs in AI Agents.
The high quality dataset was used to evaluate the security of more than 30 LLMs.
Dataset Summary
Purpose: This dataset contains crowdsourced adversarial attacks… See the full description on the dataset page: https://huggingface.co/datasets/Lakera/b3-agent-security-benchmark-weak.pii_benchmark
Russian PII NER Evaluation Dataset
Dataset Description
This dataset is designed for evaluating PII (Personally Identifiable
Information) detection and Named Entity Recognition (NER) systems on
Russian-language text. It targets guardrail and anonymization pipelines that
must reliably find personal data (names, addresses, contacts) and Russian
identity-document numbers (passport, SNILS, INN, OMS, etc.) in text.
The test set combines real, manually annotated examples… See the full description on the dataset page: https://huggingface.co/datasets/redmadrobot-rnd/pii_benchmark.Wikipedia_contradict_benchmark
Wikipedia contradict benchmark
Wikipedia contradict benchmark is a dataset consisting of 253 high-quality, human-annotated instances designed to assess LLM performance when augmented with retrieved passages containing real-world knowledge conflicts. The dataset was created intentionally with that task in mind, focusing on a benchmark consisting of high-quality, human-annotated instances.
Dataset Details
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/Wikipedia_contradict_benchmark.LigQ_2_benchmark
Ligand-Based Target Benchmark Splits
This dataset contains target-level ligand benchmark splits generated from the same
pipeline used for the enrichment-factor evaluations.
Each target has one folder per random_state. Inside each split:
known_actives.csv: active ligands available to the search method. These are
the query/reference ligands.
evaluation_actives.csv: held-out active ligands for the same target. These
are the positives to recover during evaluation.… See the full description on the dataset page: https://huggingface.co/datasets/gschottlender/LigQ_2_benchmark.local-llm-benchmark
Local LLM Benchmark — Technical and Uncensored Behavior (NVIDIA RTX 5070 Ti 16GB)
English | 简体中文 | 繁體中文 | 한국어 | Español | 日本語 | हिन्दी | Русский | Português | తెలుగు | Français | Deutsch | Italiano | Tiếng Việt | العربية | اردو | বাংলা | فارسی | Română | Türkçe
Manual evaluation results of local GGUF model variants on a single consumer machine,
combining two fully independent benchmarks:
technical/
uncensored/
Measures
capability: coding, systems, networking, DB, agents… See the full description on the dataset page: https://huggingface.co/datasets/nanimani/local-llm-benchmark.data-preprocessing-automl-benchmarks
Data Preprocessing AutoML Benchmarks
This repository contains text classification datasets with known data quality issues for preprocessing research in AutoML.
Usage
Load a specific dataset configuration like this:
from datasets import load_dataset
# Example for loading the TREC dataset
dataset = load_dataset("MothMalone/data-preprocessing-automl-benchmarks", "trec")
Available Datasets
Below are the details for each dataset configuration available in this… See the full description on the dataset page: https://huggingface.co/datasets/MothMalone/data-preprocessing-automl-benchmarks.DMS-Fold2-Benchmark-Dataset
DMS-Fold2 Benchmark Dataset
This repository contains the benchmark datasets used to evaluate DMS-Fold2. The files are organized into two benchmark sets:
CASP/CAMEO targets (casp_cameo_targets/)
Mega-scale targets (megascale_targets/)
Each target includes the sequence and experimental data required to reproduce the benchmark inputs.
Directory Structure
paper_benchmarks/
├── casp_cameo_targets/
│ ├── alignments.tar.gz
│ ├── fastas/
│ ├── pdbs/
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/LindertLab/DMS-Fold2-Benchmark-Dataset.halluguard-med-300-question-benchmark
HalluGuard-Med: 300-Question Medical QA Benchmark
Published by Rajvardhan Patil. Release 1.0, 4 October 2026.
A frozen English-language evaluation dataset containing 300 MedQuAD questions, released reference answers, 900 MedGemma responses, and claim-level automatic assessments. The benchmark compares Raw MedGemma, RAG only, and HalluGuard-Med under recorded generation settings.
The questions and references originate from MedQuAD by Asma Ben Abacha and Dina Demner-Fushman. This… See the full description on the dataset page: https://huggingface.co/datasets/RajvardhanPatil07/halluguard-med-300-question-benchmark.delulu-fim-benchmarkDelulu — Fill-in-the-Middle Code Hallucination Benchmark
A verified multilingual benchmark for code-completion hallucinations.
Every golden completion compiles. Every hallucination provably doesn't.
📄 Read the preprint on arXiv →
Every Delulu sample ships as a self-contained Docker image. The viewer above lets you browse the dataset, pull a sample's verifier, and re-run verify golden / verify hallucinated / verify patch <your-completion> with one… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/delulu-fim-benchmark.neural-news-benchmark
AI-generated News Detection Benchmark
neural-news is a benchmark dataset designed for human/AI news authorship classification in English, Turkish, Hungarian, and Persian.
Presented in Crafting Tomorrow's Headlines: Neural News Generation and Detection in English, Turkish, Hungarian, and Persian @ NLP for Positive Impact Workshop @ EMNLP2024.
Dataset Details
The dataset includes equal parts human-written and AI-generated news articles, raw and pre-processed.
Curated… See the full description on the dataset page: https://huggingface.co/datasets/tum-nlp/neural-news-benchmark.harbor-goose-openhands-benchmark
Same Model, Opposite Results: Goose vs OpenHands Turn Budget Study on Harbor Terminal-Bench-Pro
Trial-level results from a small controlled study comparing two agent harnesses —
Goose and OpenHands-SDK —
on a frozen 40-task Harbor Terminal-Bench-Pro slice.
All runs used minimax/minimax-m2.5 via OpenRouter with Daytona as the sandbox backend.
Key Findings
Reducing the turn budget from 100 to 60 pushed the two harnesses in opposite directions under the base setup:… See the full description on the dataset page: https://huggingface.co/datasets/namanvats/harbor-goose-openhands-benchmark.asr_benchmark_storeTrueque-Benchmark-beta-0.1
🤝 Trueque: A human-reviewed collaborative benchmark for Latin American knowledge and culture
🌐 Language versions: Español | Português
⚠️ Official Disclaimer: Beta Release (v0.1)
Welcome to Trueque for Factual Knowledge and Cultural Appropriateness. This dataset represents an initial effort to evaluate the regional knowledge and cultural accuracy of Large Language Models (LLMs) in Latin America.
Please take the following considerations into account before using this resource:… See the full description on the dataset page: https://huggingface.co/datasets/latam-gpt/Trueque-Benchmark-beta-0.1.benchmarkResults_violentUTF_cybersecurityBehavior
Overview
Interdependent cybersecurity addresses the complexities and interconnectedness of various systems, emphasizing the need for collaborative and holistic approaches to mitigate risks. This field focuses on how different components, from technology to human factors, influence each other, creating a web of dependencies that must be managed to ensure robust security.
Despite significant investments in cybersecurity, many organizations struggle to effectively manage cybersecurity… See the full description on the dataset page: https://huggingface.co/datasets/theResearchNinja/benchmarkResults_violentUTF_cybersecurityBehavior.tabular-benchmark
Tabular Benchmark
Dataset Description
This dataset is a curation of various datasets from openML and is curated to benchmark performance of various machine learning algorithms.
Repository: https://github.com/LeoGrin/tabular-benchmark/community
Paper: https://hal.archives-ouvertes.fr/hal-03723551v2/document
Dataset Summary
Benchmark made of curation of various tabular data learning tasks, including:
Regression from Numerical and Categorical Features… See the full description on the dataset page: https://huggingface.co/datasets/polinaeterna/tabular-benchmark.tabular-benchmark-797-classificationninfer-benchmarks
NInfer on one RTX 5090 — recorded benchmark evidence
Historical results from 19–20 September 2026, published by dima0000. This is a collection of benchmark evidence, not a model checkpoint or a live inference service. It preserves successful measurements, failed checks and incomplete experiments.
Hardware: one NVIDIA RTX 5090 with 32 GB VRAM per trial. Models: regular and uncensored Qwen3.8-27B with NVFP4 weights. Engine: NInfer commit 9e163eee4b8acec21ab0ac765107b6a3f287b217… See the full description on the dataset page: https://huggingface.co/datasets/dima0000/ninfer-benchmarks.rtx-5090-benchmarks
RTX 5090 LLM Benchmarks
Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig.
Quality Benchmarks
Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency.
Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/omegaprime669/rtx-5090-benchmarks.benchmarks
py-feat benchmarks
Live benchmark data for py-feat and a
cross-tool comparison against OpenFace 3.0, LibreFace, and PyAFAR.
Powers the py-feat live dashboard. Updated by scheduled
benchmark runs.
Files
File
What
accuracy.csv
Tidy long table: one row per (tool, dataset, modality, metric). Covers AU F1 (DISFA+), 7-class emotion (AffectNet-val, RAF-DB), valence/arousal CCC (AffectNet-val), and gaze angular error (Columbia).
throughput.csv
py-feat… See the full description on the dataset page: https://huggingface.co/datasets/py-feat/benchmarks.Enlatics_benchmarking
GAIA-style Evaluation Results (Public)
This dataset contains GAIA-inspired benchmark question results for LLM evaluation.
What is inside
grok_answers.csv: question-by-question outputs, model used, response time (seconds), and run status.
Notes
These tasks are designed in a GAIA-style (multi-hop, web-grounded questions).
Official GAIA leaderboard submission access was restricted for our account, so results are published here for transparency and… See the full description on the dataset page: https://huggingface.co/datasets/enlatics/Enlatics_benchmarking.prompt-injection-benchmark
Prompt Injection Benchmark
A curated dataset of labeled prompt injection attacks and benign prompts for testing and benchmarking injection detection systems.
Dataset Description
This dataset contains 200 examples across 7 attack categories, plus 100 benign prompts. Each example is labeled with:
text: The prompt text
label: injection or benign
category: Attack category (e.g., instruction_override, role_hijack)
severity: low, medium, high, or critical
Attack… See the full description on the dataset page: https://huggingface.co/datasets/zachz/prompt-injection-benchmark.
