datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Nemotron-ClimbMix
ClimbMix Dataset
🚀 Creating the highest-quality pre-training datasets for LLMs 🌟
📄 PAPER
🤗 CLIMBLAB
🤗 CLIMBMIX
🏠 HOMEPAGE
Figure 1: Continuously training a 1B model yields a 2.0% improvement over Llama-3.2-1B, demonstrating a more efficient scaling trend compared to prior models.
Figure 2: Pre-training a 1B model from scratch on ClimbMix shows better scaling effects than training on other datasets.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-ClimbMix.ClimbMixClimbMix is a high-quality pre-training corpus released by NVIDIA. Here is the description:
ClimbMix is a compact yet powerful 400-billion-token dataset designed for efficient pre-training that delivers superior performance under an equal token budget. It was introduced in this paper.
We proposed a new algorithm to filter and mix the dataset. First, we grouped the data into 1,000 groups based on topic information. Then we applied two classifiers: one to detect advertisements and another to… See the full description on the dataset page: https://huggingface.co/datasets/OptimalScale/ClimbMix.Real-UI-Clickboxes
RUC: Real UI Clickboxes
Click carefully, even when the page is trying to trick you! 👀
Official Hugging Face release for RUC: Real UI Clickboxes, the dataset accompanying our ACL 2026 paper Don't Click That: Teaching Web Agents to Resist Deceptive Interfaces on deceptive UI understanding for web agents.
ACL Anthology: https://aclanthology.org/2026.acl-long.310/
PDF: https://aclanthology.org/2026.acl-long.310.pdf
DOI: https://doi.org/10.18653/v1/2026.acl-long.310… See the full description on the dataset page: https://huggingface.co/datasets/DUDE-Framework/Real-UI-Clickboxes.climbing-holds
[!IMPORTANT]
This dataset is in construction. The current files are raw scans intended for establishing the structure.
Using them? Help us clean them up or identify the brands by consulting the CONTRIBUTING.md guide.
GUI for contributions
https://setrsoft.github.io/holds-dataset-hub/
Or send your files here
Climbing Holds 3D dataset (SetRsoft)
📋 Project Overview
This dataset is a community-driven open-source dataset of 3D-scanned climbing holds… See the full description on the dataset page: https://huggingface.co/datasets/setrsoft/climbing-holds.sih-lidar-clipClinSeek-Bench
ClinSeek-Bench
ClinSeek-Bench is the evaluation suite introduced in
ClinSeekAgent: Automating Multimodal Evidence Seeking for Agentic Clinical
Reasoning. It evaluates clinical reasoning
under two paired settings with the same task definitions and answer labels:
Curated Input: the model answers from the evidence package provided by
the source benchmark.
Automated Evidence-Seeking: the curated context is removed, and the model
must retrieve evidence from raw clinical data using… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/ClinSeek-Bench.CLI-Bench
CLI-Bench: Benchmarking AI Agents on Command-Line Tool Orchestration
Abstract
CLI-Bench is an evaluation benchmark for measuring AI agents' ability to learn and use command-line interface (CLI) tools to complete real-world tasks. Unlike existing benchmarks that test general coding ability or narrow tool-use scenarios, CLI-Bench evaluates tool-agnostic CLI orchestration -- the capacity to read tool documentation, plan multi-step workflows, execute commands… See the full description on the dataset page: https://huggingface.co/datasets/hiklikai/CLI-Bench.ai-climate-compute-2026
Ai Climate Compute 2026
Part of the LEGION Intelligence dataset collection.
Provider: LEGION Systems
Access: Requires approval — submit request below
Usage
from datasets import load_dataset
dataset = load_dataset("gemmozero/ai-climate-compute-2026")
API Access
Real-time access via LEGION API:
curl https://api.legion-api.com/incidents
API Docs · Pro Access €29/mo
License
CC BY-NC 4.0 — Research and non-commercial use only.
Commercial… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-climate-compute-2026.mteb-nl-sarcastic-headlines This dataset contains news headlines from a satirical news website (Speld.nl) and a regular news website that are annotated with binary sarcasm labels (1 indicating sarcasm, 0 indicating non-sarcasm). All headlines from Speld.nl are annotated as sarcastic, whereas all headlines from nu.nl are not.
Citation Information
If you find our paper, benchmark or models helpful, please consider cite as follows:
@misc{banar2025mtebnle5nlembeddingbenchmark,
title={MTEB-NL and E5-NL:… See the full description on the dataset page: https://huggingface.co/datasets/clips/mteb-nl-sarcastic-headlines.pentabrid-reproducibility
Pentabrid 27B: reproducibility package
Everything required to recompute the results of a controlled evaluation of fine-tuning
configurations for medical question answering. Openly available with no access
restrictions.
Contents
Path
Description
per_item/medxpertqa_*.jsonl
Per-item predictions for all six checkpoints on 2,450 MedXpertQA-Text items. Fields: id, gold, extracted_answer, correct, explicit_marker_present, n_markers, response_chars… See the full description on the dataset page: https://huggingface.co/datasets/Clinical-Reasoning-Hub/pentabrid-reproducibility.earth-love-united-climate-knowledge
🌍 Earth Love United Climate Knowledge Dataset
The most comprehensive open climate science knowledge dataset.
10,128 text chunks + 124 structured facts + 4.54B year geological memory + 10 tipping points.
Built to power GAIA — an AI that embodies the living consciousness of Earth.
Dataset Overview
This dataset gives an AI system authoritative, sourced knowledge about climate change,
carbon, Earth science, and solutions. It has four layers:
Layer 1: Text Knowledge… See the full description on the dataset page: https://huggingface.co/datasets/ego0op/earth-love-united-climate-knowledge.tanzania-clinical-evidence
Tanzania Clinical Evidence Dataset — draft v2
This repository contains a traceable evidence corpus extracted from Tanzania health documents and the STG/NEMLIT 7th Edition 2026 collection. It provides source-text chunks for continued pretraining, source-grounded instruction drafts, constrained RLVR tasks, and structured NEMLIT tables.
Status
This is a draft research dataset. All examples remain pending qualified clinical review and source-family split assignment.… See the full description on the dataset page: https://huggingface.co/datasets/Japhari/tanzania-clinical-evidence.qwen3-4b-instruct-bestatklyra-climbmix-30b
Details
A 30B-token slice of NVIDIA ClimbMix
(via the shuffled karpathy/climbmix-400b-shuffle),
pre-tokenized and prepared as pretraining data for the Lyra project.
Dataset
Size: 30 billion tokens
Format: ArrayRecord
Tokenizer: o200k_harmony — tiktoken o200k_base plus the gpt-oss special tokens (vocab 201,088)
Shards: 100 ArrayRecord files (group_size:1), exactly 300M tokens each
No BOS token; documents are delimited by EOS only
cli-commands-explained
Overview
This dataset is a collection of 16,098 command line instructions sourced from Commandlinefu and Cheatsheets. It includes an array of commands, each with an id, title, description, date, url to source, author, votes, and flag indicating if the description is AI generated. The descriptions are primarily authored by the original contributors, for entries where descriptions were absent, they have been generated using NeuralBeagle14-7B. Out of the total entries, 10,039… See the full description on the dataset page: https://huggingface.co/datasets/b-mc2/cli-commands-explained.fitllm-fit-census
Local LLM Fit Census v1 — 2026-09-13
10,530 verdicts: 30 models × 93 devices (36 GPUs + 57 Mac configs) × per-platform quant tiers.
Each row is generated by fitllm-engine from architecture inputs pinned to official config.json files. Runtime and OS reserves remain documented estimates. Reproduce it yourself: npm run census.
Assumptions: context = min(8K, model max) · KV cache F16 · platform reserve/headroom per engine. Interactive per-combo pages: fitllm.run/can-i-run.… See the full description on the dataset page: https://huggingface.co/datasets/click6067/fitllm-fit-census.single-click_bench
Single-Click Benchmark for Web Interaction
This benchmark defines a minimal web interaction task that requires only a single click to complete.
Each instance includes two task formulations (simplified and human-like), a pre-saved HTML file for obtaining screenshots or metadata, and target annotations with the element’s bounding box and XPath for evaluation.
The dataset enables systematic evaluation of web agents’ capabilities such as visual grounding, task understanding, and action… See the full description on the dataset page: https://huggingface.co/datasets/alexandrayakovleva/single-click_bench.atlaspi-historical-geography
AtlasPI — Historical Geography Dataset
1,006 historical geopolitical entities · 643 events · 55 periods · 104 dynasty chains · 252 cities · 41 trade routes
The first open dataset specifically designed for AI agents working on
historical geography questions. Apache 2.0 licensed. Includes real GeoJSON
boundaries from academic sources, not placeholder polygons.
Temporal range: 4500 BCE → 2024 CE
Geographic coverage: all inhabited continents (Asia 31%, Africa 18%,
Americas 17%… See the full description on the dataset page: https://huggingface.co/datasets/clirim911/atlaspi-historical-geography.ClimbMix-sample
Unofficial NVIDIA Nemotron-ClimbMix (Subsampled)
This dataset is a curated, subsampled version of OptimalScale/ClimbMix, which itself is a detokenized version of NVIDIA's official pretraining dataset, nvidia/Nemotron-ClimbMix.
It is designed for researchers and developers looking for a smaller, well-shuffled slice of the Nemotron pretraining data for quick experimentation, testing, or ablation studies.
Processing Method
To create this streamlined version, the… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/ClimbMix-sample.climate-fever-v2
ClimateFEVER.v2
An MTEB dataset
Massive Text Embedding Benchmark
CLIMATE-FEVER is a dataset following the FEVER methodology, containing 1,535 real-world climate change claims. This updated version addresses corpus mismatches and qrel inconsistencies in MTEB, restoring labels while refining corpus-query alignment for better accuracy.
Task category
t2t
Domains
Academic, Written
Reference
https://www.sustainablefinance.uzh.ch/en/research/climate-fever.html… See the full description on the dataset page: https://huggingface.co/datasets/mteb/climate-fever-v2.SERA-KimiK3-Django-SWEAgent-Cliff32k-T1
SERA Kimi-K3 Django SWE-Agent — Cliff-chunked T1 (first rollout)
572 training records built from 210 Kimi-K3 SWE-agent trajectories on
Django, split to fit a 32,768-token context with
CliffCompaction instead of being truncated.
Why chunked
A 100+ step agent rollout does not fit a 32k training window — 76% of the
source T1 trajectories exceed it. Truncating them throws away most of the
supervision, and trains the model on a context format it never sees at… See the full description on the dataset page: https://huggingface.co/datasets/thientrangngv/SERA-KimiK3-Django-SWEAgent-Cliff32k-T1.spectralbio-clinvar
SpectralBio Dataset Card
Public Hierarchy
SpectralBio separates manuscript-facing scientific centrality from frozen executable replay centrality.
Flagship scientific result: BRCA2 covariance-aware augmentation against a stronger five-model ESM-1v baseline
Validation anchor: TP53 is the only frozen public canonical replay surface
Breadth surface: support-ranked top-25 feasible panel derived from the 15,752-gene ClinVar scan
Boundary surfaces: protocol sweep and BRCA1… See the full description on the dataset page: https://huggingface.co/datasets/DaviBonetto/spectralbio-clinvar.stride-preproc-climbmix
STRIDE: Preprocessed ClimbMix
Tokenized ClimbMix sequences used for STRIDE's pretraining-data attribution experiments. There is one training pool per released nanochat depth (d12, d16, d20, d24) and one held-out test set shared across depths. These are the corpora the released nanochat operators attribute over, so STRIDE's column indices line up with the lines of these files.
Files
File
Sequences
Size
Contents
climbmix_train_d12.jsonl
1,317,003
3.8 GB
training… See the full description on the dataset page: https://huggingface.co/datasets/CausalNLP/stride-preproc-climbmix.SERA-KimiK3-Django-SWEAgent-Cliff32k-T2
SERA Kimi-K3 Django SWE-Agent — Cliff-chunked T2 (second rollout)
227 training records built from 137 Kimi-K3 SWE-agent trajectories on
Django, split to fit a 32,768-token context with
CliffCompaction instead of being truncated.
Why chunked
A 100+ step agent rollout does not fit a 32k training window — 27% of the
source T2 trajectories exceed it. Truncating them throws away most of the
supervision, and trains the model on a context format it never sees at… See the full description on the dataset page: https://huggingface.co/datasets/thientrangngv/SERA-KimiK3-Django-SWEAgent-Cliff32k-T2.finject
FInject Dataset Card
FInject is a financial unanswerability benchmark built by transforming answerable financial reasoning problems into controlled unanswerable variants. Each row preserves the original question and pairs an answerable original context with a perturbed context that is no longer sufficient to support a unique answer.
Dataset Summary
Seed source: 78 answerable hard problems from FinanceReasoning.
Final release size: 426 unanswerable variants.… See the full description on the dataset page: https://huggingface.co/datasets/pnu-clink/finject.improved_aesthetics_6.5plus_clip_retrievalclickhouse-server-imagesurgeon-tested-clinical-ai-benchmark
Surgeon-Tested Clinical AI Benchmark (TH-CAB v1.1)
An independent, reproducible evaluation of large language models (LLMs) on real, de-identified cancer cases — scored by a practicing surgeon item-by-item against current clinical guidelines.
Homepage & full leaderboard: https://tanhaosheng.asia
Methodology (citable authority, TH-CAB v1.1): https://tanhaosheng.asia/methodology/
Open data layer: https://tanhaosheng.asia/data/
This is a benchmark / research dataset, not clinical… See the full description on the dataset page: https://huggingface.co/datasets/tanhaosheng/surgeon-tested-clinical-ai-benchmark.Product_reviews_click_stream_DataaClinSeekAgent_RL
ClinSeekAgent_RL
RL-ready prompt sets derived from
Letian2003/DeepMed_trajectory.
The source release ships agent rollout trajectories: every question appears 4–5
times, once per sampled run, each carrying the full messages transcript and that
run's outcome. That shape is built for SFT and trajectory analysis. This derivative
strips it back to what an RL loop actually needs — the prompt and its ground-truth
label, once per question — so the policy generates its own trajectories… See the full description on the dataset page: https://huggingface.co/datasets/Chtholly17/ClinSeekAgent_RL.
