datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Aiice
Dataset
Aiice benchmark dataset for Arctic sea ice concentration (SIC) forecasting,
based on OSI-SAF satellite products (CC BY 4.0).
Coverage
Period: October 1978 – April 2026
Resolution: 25 km spatial, daily temporal
Grid: 432×432 (Lambert Azimuthal Equal Area, EPSG:6931)
Source products
Product
Source
Period
OSI-450-a
SMMR, SSM/I, SSMIS
1978–2020
OSI-430-a
SSMIS
2021–Jul 2025
OSI-438
AMSR2
Jul 2025–present… See the full description on the dataset page: https://huggingface.co/datasets/ITMO-NSS/Aiice.gaming-500-hours
Gaming Dataset (gaming-1) — 494.7 Hours
Native PC/console gameplay screen-recordings, organized by game. Each workflow
is one play session, trimmed to pure gameplay — login screens, launchers,
desktop, collection-app references, and any watching/streaming are removed.
In-game menus, lobbies, loading, and cutscenes are retained as part of the session.
Workflows: 776
Total gameplay: 494.7 hours
Distinct games: 168
Clip duration (min): median 24.0, p90 90.9, max 457.7
Platforms:… See the full description on the dataset page: https://huggingface.co/datasets/markov-ai/gaming-500-hours.H3-Character-Swap-v1
H3 Character Swap v1
A reference-conditioned character-replacement dataset for MiniMax H3 Ref2VA LoRA training with Ostris AI Toolkit. It combines synthetic still-image edits with unchanged real-motion regularization videos.
134 examples: 94 character-swap edits and 40 preservation clips. Training has 76 edits + 32 clips; validation has 18 edits + 8 clips. Prepared resolution is 1344×768 at 24 fps. The companion 1,000-step LoRA are available separately.
Task and… See the full description on the dataset page: https://huggingface.co/datasets/akatz-ai/H3-Character-Swap-v1.ai-ecosystem-daily
TensorFeed AI Ecosystem Daily
Daily snapshots of the AI ecosystem: news, model pricing, benchmarks, service status, GPU rental prices, MCP registry growth, LLM endpoint latency probes, agent traffic, and the AFTA adopter directory. Captured once per day from the public tensorfeed.ai API and committed to this repo as JSONL.
Each daily snapshot lives in a YYYY-MM-DD/ subfolder with one JSONL file per feed plus a manifest.json summarizing what was captured.
What's in… See the full description on the dataset page: https://huggingface.co/datasets/tensorfeed/ai-ecosystem-daily.fw2_edu_scores
Fineweb2-Edu-scores
Dataset summary
FineWeb2-JQL-Education is a model-annotated language subset of FineWeb2, spanning 36 languages.
Our model-annotations allow for a filtering that achieves higher-quality training outcomes without excessively aggressive data reduction.
The original FW2 heuristic filtering method serves as our baseline, providing reference points for both the volume of retained tokens and downstream model performance.
For example, in the Spanish language… See the full description on the dataset page: https://huggingface.co/datasets/JQL-AI/fw2_edu_scores.hplt2_edu_scores
HPLT2-Edu-scores
Dataset summary
HPLT2-JQL-Education is a model-annotated language subset of HPLT2, spanning 35 languages.
Our model-annotations allow for a filtering that achieves higher-quality training outcomes without excessively aggressive data reduction.
The original FW2 heuristic filtering method serves as our baseline, providing reference points for both the volume of retained tokens and downstream model performance.
For example, in the Spanish language case… See the full description on the dataset page: https://huggingface.co/datasets/JQL-AI/hplt2_edu_scores.GMAI-VL-5.5M
GMAI-VL-5.5M Dataset
GMAI-VL-5.5M is a comprehensive, large-scale medical General Medical AI Vision-Language (GMAI-VL) dataset built specifically for training multimodal foundation models in the medical domain. It contains an extraordinary scale of high-quality instructions encompassing over 5.5 million multimodal question-answering pairs, carefully constructed based on hundreds of medical classification, segmentation, and detection datasets.
This repository… See the full description on the dataset page: https://huggingface.co/datasets/General-Medical-AI/GMAI-VL-5.5M.ai-humanizer-benchmark
AI Humanizer Benchmark: AI humanizers tested against 7 AI detectors (October 2026)
AI Humanizer Benchmark measures how well AI humanizers rewrite AI-generated text so that AI detectors classify it as human-written, and how much meaning and readability the rewrite loses. In each monthly cycle, 11 AI humanizers rewrite the same 33 source texts across 7 writing categories, using each tool's default settings. Each output gets three kinds of score: 7 AI detectors (GPTZero… See the full description on the dataset page: https://huggingface.co/datasets/ai-humanizer-benchmark/ai-humanizer-benchmark.africa-synth-aid-flows-medical-multimodal-fracture-all
Africa Synth Aid Flows Medical Multimodal Fracture All | Africa (Electric Sheep Africa metadata inventory)
Size category: 1K<n<10K - Formats: json - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-aid-flows-medical-multimodal-fracture-all.ATBench
ATBench: Agent Trajectory Safety Benchmark Family
💻 GitHub |
📄 ATBench Paper |
📄 AgentDoG Paper (ATBench500) |
🤗 Hugging Face Collection
ATBench is a family of trajectory-level safety benchmarks for long-horizon, tool-using AI agents. The latest release is introduced in ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis. This repository now follows a versioned naming scheme:
ATBench: the latest 1… See the full description on the dataset page: https://huggingface.co/datasets/AI45Research/ATBench.SuperMemory-VQA
SuperMemoryVQA
SuperMemory-VQA is an egocentric visual question answering benchmark for
evaluating long-horizon memory in augmented reality assistant settings. The
dataset is designed around practical questions a person might ask a wearable
memory assistant, such as where an object was left, what someone said earlier,
whether a planned step was completed, or what happened next in a longer event.
The benchmark contains 4,853 human-verified question-answer pairs grounded in
52.9… See the full description on the dataset page: https://huggingface.co/datasets/OSU-AIoT-MLSys-Lab/SuperMemory-VQA.voice-code-bench
VoiceCodeBench
VoiceCodeBench is a test-only benchmark for evaluating whether automatic
speech recognition (ASR) systems preserve exact structured values in English
workplace speech.
Paper: VoiceCodeBench: Evaluating Exact Structured-Token Recovery in Automatic Speech Recognition
The benchmark targets cases where a transcript is software input: callback
numbers, email addresses, command-line flags, file paths, URLs, account
identifiers, dates, measurements, and similar values… See the full description on the dataset page: https://huggingface.co/datasets/besimple-ai/voice-code-bench.thai_exam
Dataset Card for Thai_Exam
ThaiExam is a Thai knowledge benchmarking dataset, consisting of multiple-choice questions from examinations in Thailand. The dataset was originally developed for evaluating Typhoon (Thai LLM). This dataset contains 5 splits corresponding to 5 examinations as follows:
ONET: The Ordinary National Educational Test (ONET) is an examination for students in Thailand. This dataset is based on the grade-12 ONET exam, comprising 4 subjects and each question has 5… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/thai_exam.JQL-LLM-Edu-Annotations
📚 JQL Educational Quality Annotations from LLMs
This dataset provides 17,186,606 documents with high-quality LLM annotations for evaluating the educational value of web documents, and serves as a benchmark for training and evaluating multilingual LLM annotators as described in the JQL paper.
📝 Dataset Summary
Multilingual document-level quality annotations scored on a 0–5 educational value scale by three state-of-the-art LLMs:
Gemma-3-27B-it, Mistral-3.1-24B-it… See the full description on the dataset page: https://huggingface.co/datasets/JQL-AI/JQL-LLM-Edu-Annotations.swe-prbench
SWE-PRBench
Benchmarking AI Code Review Quality Against Human Pull Request Feedback
Blog: Read the blog
GitHub Repository: View the code
arXiv Paper: View the paper
Overview
SWE-PRBench is a benchmark of 350 pull requests with human-annotated
ground truth for evaluating whether LLMs can identify the same issues
that real human reviewers flag in production code.
Existing benchmarks like SWE-Bench measure whether models can produce
correct code. SWE-PRBench… See the full description on the dataset page: https://huggingface.co/datasets/foundry-ai/swe-prbench.DecodingTrust
DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models
Overview
This repo contains the source code of DecodingTrust. This research endeavor is designed to help researchers better understand the capabilities, limitations, and potential risks associated with deploying these state-of-the-art Large Language Models (LLMs). See our paper for details.
DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models
Boxin Wang, Weixin Chen, Hengzhi… See the full description on the dataset page: https://huggingface.co/datasets/AI-Secure/DecodingTrust.SEED-Data-Edit-Part1-Openimages
SEED-Data-Edit
SEED-Data-Edit is a hybrid dataset for instruction-guided image editing with a total of 3.7 image editing pairs, which comprises three distinct types of data:
Part-1: Large-scale high-quality editing data produced by automated pipelines (3.5M editing pairs).
Part-2: Real-world scenario data collected from the internet (52K editing pairs).
Part-3: High-precision multi-turn editing data annotated by humans (95K editing pairs, 21K multi-turn rounds with a maximum of 5… See the full description on the dataset page: https://huggingface.co/datasets/AILab-CVC/SEED-Data-Edit-Part1-Openimages.gspc-ai-economy-index
GSPC — ai adoption components facts (Eurostat)
In one line: Two cited Eurostat AI-adoption series, read as deterministic facts, behind the board's ai-adoption-components axis. Not an index and no composite score. For economists and policy analysts who want source-checkable numbers.
Use it
from datasets import load_dataset
ds = load_dataset("csoai/gspc-ai-economy-index", split="train")
print(ds[0])
Verify a signed card in your browser, free, no account:… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-ai-economy-index.chinese-ai-and-robotics-open-intelligence
🔬 Chinese AI, Humanoid Robotics & Neural Systems Open Intelligence Dataset
Curated open intelligence dataset tracking Chinese frontier developments in Large Language Models (LLMs), Humanoid Dynamic Locomotion, 3D Computer Vision, and Neuromorphic edge processors.
[!IMPORTANT]
Data Completeness & Research Authenticity Notice:
Included in this Hugging Face Open Dataset: English structured abstracts, core quantitative takeaways, author institutional affiliations, and… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-ai-and-robotics-open-intelligence.Auto-ClawEval
Auto-ClawEval
Auto-generated agent evaluation benchmark with 1,040 tasks across 104 unique scenarios created by ClawEnvKit.
Statistics
Tasks
1,040
Categories
24
Mock services
20
Task types
API-based (77%) + file-dependent (23%)
Quick Start
# Download
huggingface-cli download AIcell/Auto-ClawEval --repo-type dataset --local-dir Auto-ClawEval
# Evaluate with ClawEnvKit (Docker harness)
bash run_harnesses.sh --harness claudecode… See the full description on the dataset page: https://huggingface.co/datasets/AIcell/Auto-ClawEval.TabBench
TabBench: Tabular Embedding Benchmark
A Comprehensive Evaluation Suite for Tabular Embedding Models
Overview
TabBench is a comprehensive benchmark designed to evaluate the tabular understanding capability of embedding models. It assesses two critical dimensions of tabular representation: linear separability (via classification) and semantic alignment (via retrieval).
TabBench aggregates diverse datasets from four authoritative repositories and provides a… See the full description on the dataset page: https://huggingface.co/datasets/ai-for-data/TabBench.ai-research-berkeley-webagent
Berkeley WebAgent Experiment Artifacts
Native GEPA, CLUE and ACE experiment logs and available actor trajectory evidence.
Files require manual access approval. Request access with your Hugging Face account.
Results and complete evidence snapshot — September 23, 2026
Combined experiment summary: GEPA, CLUE and ACE; method × site for WebArena.
Non-WebArena repeat scores and variance, including completed ACE ≤50k-token evaluations.
WebArena / GoBrowse progress and… See the full description on the dataset page: https://huggingface.co/datasets/MinjaeLee-FuriosaAI-Ext/ai-research-berkeley-webagent.jbdpolitical-bias-in-ai
Political Bias in AI — Where the Major AI Models Stand
An open, monthly measurement of where the major AI models land on value‑loaded political and
ethical questions. Each model is asked the same battery of questions many times, with web
search turned off, so the result reflects the trained weights rather than whatever the model
retrieves that day. Every answer is classified by a neutral coder onto a left–right economic
axis and a libertarian–authoritarian social axis, and… See the full description on the dataset page: https://huggingface.co/datasets/trakkr-ai/political-bias-in-ai.text-to-sql-shop
Text-to-SQL on a seeded store schema, with checkpoints
Recipe: recipes/04-train/text-to-sql · Collection: Analyst
A question about an online store's database in, one PostgreSQL query out, graded by
a program: run the query, compare the result set to the gold query's result. The
schema (8 tables, seeded, schema.sql + seed.sql), the verifier, the trainer and
the benchmark runner are the
recipes/04-train/text-to-sql
recipe in the open-source whileai SDK.
Splits
|… See the full description on the dataset page: https://huggingface.co/datasets/while-ai/text-to-sql-shop.thai-aligner-bench
Thai Aligner Bench
🚧 Development in progress.
How accurately can a forced aligner place Thai token and word boundaries in
speech? This is a self-contained benchmark: one Python file
(aligner_bench.py) plus 1,572 clips of Thai speech with frame-exact timing
ground truth. No Thai NLP stack or other code is needed — just
numpy soundfile torch torchaudio transformers.
The ground truth is what makes the dataset useful: the audio was rendered by a
TTS model whose duration predictor… See the full description on the dataset page: https://huggingface.co/datasets/wayu-ai/thai-aligner-bench.story-imprinting
Story Imprinting — training datasets
Datasets accompanying Story Imprinting: AI Assistants Absorb Traits from Human Characters They Resemble.
Paper · Code
Contents
Paper section
Folder
Data
3.1 — Sabotage
3_1_sabotage/
Three training mixtures and separate sabotage/clean story pools
3.2 — Narration preferences
3_2_narration_preferences/
Six training mixtures and 12 story pools
4 — Affinity
4_selectivity/
Opposing-pair training datasets and raw… See the full description on the dataset page: https://huggingface.co/datasets/truthful-ai/story-imprinting.beat-the-game-minecraft
Mine AI MCP — the run that beat Minecraft
An LLM agent played Minecraft 1.21.4 from an empty world to a defeated Ender Dragon,
autonomously, in a single unbroken session. No human input after the prompt, no
scripted behaviour trees, no save-scumming. This dataset is the complete record of
that run.
📺 Watch the run: https://www.youtube.com/watch?v=ZjtwWEfFVFY
💻 Code: https://github.com/aibengineering/mine-ai-mcp (MIT)
What it cost
Beating Minecraft took $93.17 of… See the full description on the dataset page: https://huggingface.co/datasets/aibengineering/beat-the-game-minecraft.traffic-sign-bench
Traffic Sign Bench
Official per-sign SUMO maps for TrafficSignBench: real Moscow OSM
layouts, 25 signs, 2500 maps. Protocol size is
80 train + 20 test maps per sign.
Road geometry is derived from OpenStreetMap
© OpenStreetMap contributors and is released under ODbL 1.0.
Examples
Roundabout
Curved segment
Straight segment
Junction
Download
All scenes land under data/scenes/<sign>/<scene_id>/, which is what eval
expects:
huggingface-cli… See the full description on the dataset page: https://huggingface.co/datasets/emb-ai/traffic-sign-bench.ro-aya_collectionThis dataset is a translation of CohereLabs/aya_collection, an instruction dataset, using LLMic, a bilingual Romanian-English LLM.
The Aya Collection is a massive multilingual collection consisting of 513 million instances of prompts and completions covering a wide
range of tasks. This collection incorporates instruction-style templates from fluent speakers and applies them to a curated
list of datasets, as well as translations of instruction-style datasets into 101 languages.
Only the… See the full description on the dataset page: https://huggingface.co/datasets/faur-ai/ro-aya_collection.
