Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nvidia /Nemotron-ClimbMix ClimbMix Dataset 🚀 Creating the highest-quality pre-training datasets for LLMs 🌟 📄 PAPER 🤗 CLIMBLAB 🤗 CLIMBMIX 🏠 HOMEPAGE Figure 1: Continuously training a 1B model yields a 2.0% improvement over Llama-3.2-1B, demonstrating a more efficient scaling trend compared to prior models. Figure 2: Pre-training a 1B model from scratch on ClimbMix shows better scaling effects than training on other datasets.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-ClimbMix.tabulartext-generation100M<n<1B131 likes6.8k downloads1y agoHugging Face02OptimalScale /ClimbMixClimbMix is a high-quality pre-training corpus released by NVIDIA. Here is the description: ClimbMix is a compact yet powerful 400-billion-token dataset designed for efficient pre-training that delivers superior performance under an equal token budget. It was introduced in this paper. We proposed a new algorithm to filter and mix the dataset. First, we grouped the data into 1,000 groups based on topic information. Then we applied two classifiers: one to detect advertisements and another to… See the full description on the dataset page: https://huggingface.co/datasets/OptimalScale/ClimbMix.tabulartext-generation100M<n<1B38 likes1.5k downloads1y agoHugging Face03DUDE-Framework /Real-UI-Clickboxes RUC: Real UI Clickboxes Click carefully, even when the page is trying to trick you! 👀 Official Hugging Face release for RUC: Real UI Clickboxes, the dataset accompanying our ACL 2026 paper Don't Click That: Teaching Web Agents to Resist Deceptive Interfaces on deceptive UI understanding for web agents. ACL Anthology: https://aclanthology.org/2026.acl-long.310/ PDF: https://aclanthology.org/2026.acl-long.310.pdf DOI: https://doi.org/10.18653/v1/2026.acl-long.310… See the full description on the dataset page: https://huggingface.co/datasets/DUDE-Framework/Real-UI-Clickboxes.imageimage-text-to-text1K<n<10K1 likes683 downloads3mo agoHugging Face04setrsoft /climbing-holds [!IMPORTANT] This dataset is in construction. The current files are raw scans intended for establishing the structure. Using them? Help us clean them up or identify the brands by consulting the CONTRIBUTING.md guide. GUI for contributions https://setrsoft.github.io/holds-dataset-hub/ Or send your files here Climbing Holds 3D dataset (SetRsoft) 📋 Project Overview This dataset is a community-driven open-source dataset of 3D-scanned climbing holds… See the full description on the dataset page: https://huggingface.co/datasets/setrsoft/climbing-holds.3dn<1K1 likes656 downloads6mo agoHugging Face05ranbyDipz /sih-lidar-cliptabularn<1K0 likes264 downloads11d agoHugging Face06UCSC-VLAA /ClinSeek-Bench ClinSeek-Bench ClinSeek-Bench is the evaluation suite introduced in ClinSeekAgent: Automating Multimodal Evidence Seeking for Agentic Clinical Reasoning. It evaluates clinical reasoning under two paired settings with the same task definitions and answer labels: Curated Input: the model answers from the evidence package provided by the source benchmark. Automated Evidence-Seeking: the curated context is removed, and the model must retrieve evidence from raw clinical data using… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/ClinSeek-Bench.tabular1K<n<10K2 likes185 downloads1mo agoHugging Face07hiklikai /CLI-Bench CLI-Bench: Benchmarking AI Agents on Command-Line Tool Orchestration Abstract CLI-Bench is an evaluation benchmark for measuring AI agents' ability to learn and use command-line interface (CLI) tools to complete real-world tasks. Unlike existing benchmarks that test general coding ability or narrow tool-use scenarios, CLI-Bench evaluates tool-agnostic CLI orchestration -- the capacity to read tool documentation, plan multi-step workflows, execute commands… See the full description on the dataset page: https://huggingface.co/datasets/hiklikai/CLI-Bench.documenttext-generationn<1K0 likes172 downloads18d agoHugging Face08gemmozero /ai-climate-compute-2026gated Ai Climate Compute 2026 Part of the LEGION Intelligence dataset collection. Provider: LEGION Systems Access: Requires approval — submit request below Usage from datasets import load_dataset dataset = load_dataset("gemmozero/ai-climate-compute-2026") API Access Real-time access via LEGION API: curl https://api.legion-api.com/incidents API Docs · Pro Access €29/mo License CC BY-NC 4.0 — Research and non-commercial use only. Commercial… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-climate-compute-2026.tabulartext-classificationn<1K0 likes160 downloads4d agoHugging Face09clips /mteb-nl-sarcastic-headlines This dataset contains news headlines from a satirical news website (Speld.nl) and a regular news website that are annotated with binary sarcasm labels (1 indicating sarcasm, 0 indicating non-sarcasm). All headlines from Speld.nl are annotated as sarcastic, whereas all headlines from nu.nl are not. Citation Information If you find our paper, benchmark or models helpful, please consider cite as follows: @misc{banar2025mtebnle5nlembeddingbenchmark, title={MTEB-NL and E5-NL:… See the full description on the dataset page: https://huggingface.co/datasets/clips/mteb-nl-sarcastic-headlines.tabular10K<n<100K0 likes154 downloads1y agoHugging Face10Clinical-Reasoning-Hub /pentabrid-reproducibility Pentabrid 27B: reproducibility package Everything required to recompute the results of a controlled evaluation of fine-tuning configurations for medical question answering. Openly available with no access restrictions. Contents Path Description per_item/medxpertqa_*.jsonl Per-item predictions for all six checkpoints on 2,450 MedXpertQA-Text items. Fields: id, gold, extracted_answer, correct, explicit_marker_present, n_markers, response_chars… See the full description on the dataset page: https://huggingface.co/datasets/Clinical-Reasoning-Hub/pentabrid-reproducibility.tabularquestion-answering10K<n<100K0 likes143 downloads22d agoHugging Face11ego0op /earth-love-united-climate-knowledge 🌍 Earth Love United Climate Knowledge Dataset The most comprehensive open climate science knowledge dataset. 10,128 text chunks + 124 structured facts + 4.54B year geological memory + 10 tipping points. Built to power GAIA — an AI that embodies the living consciousness of Earth. Dataset Overview This dataset gives an AI system authoritative, sourced knowledge about climate change, carbon, Earth science, and solutions. It has four layers: Layer 1: Text Knowledge… See the full description on the dataset page: https://huggingface.co/datasets/ego0op/earth-love-united-climate-knowledge.tabulartext-retrieval10K<n<100K2 likes134 downloads5mo agoHugging Face12Japhari /tanzania-clinical-evidence Tanzania Clinical Evidence Dataset — draft v2 This repository contains a traceable evidence corpus extracted from Tanzania health documents and the STG/NEMLIT 7th Edition 2026 collection. It provides source-text chunks for continued pretraining, source-grounded instruction drafts, constrained RLVR tasks, and structured NEMLIT tables. Status This is a draft research dataset. All examples remain pending qualified clinical review and source-family split assignment.… See the full description on the dataset page: https://huggingface.co/datasets/Japhari/tanzania-clinical-evidence.tabular10K<n<100K0 likes92 downloads2d agoHugging Face13clijo /qwen3-4b-instruct-bestatktabular100K<n<1M0 likes82 downloads4mo agoHugging Face14Ibbyml /lyra-climbmix-30b Details A 30B-token slice of NVIDIA ClimbMix (via the shuffled karpathy/climbmix-400b-shuffle), pre-tokenized and prepared as pretraining data for the Lyra project. Dataset Size: 30 billion tokens Format: ArrayRecord Tokenizer: o200k_harmony — tiktoken o200k_base plus the gpt-oss special tokens (vocab 201,088) Shards: 100 ArrayRecord files (group_size:1), exactly 300M tokens each No BOS token; documents are delimited by EOS only tabularn<1K1 likes78 downloads12d agoHugging Face15b-mc2 /cli-commands-explained Overview This dataset is a collection of 16,098 command line instructions sourced from Commandlinefu and Cheatsheets. It includes an array of commands, each with an id, title, description, date, url to source, author, votes, and flag indicating if the description is AI generated. The descriptions are primarily authored by the original contributors, for entries where descriptions were absent, they have been generated using NeuralBeagle14-7B. Out of the total entries, 10,039… See the full description on the dataset page: https://huggingface.co/datasets/b-mc2/cli-commands-explained.tabulartext-generation10K<n<100K5 likes72 downloads2y agoHugging Face16click6067 /fitllm-fit-census Local LLM Fit Census v1 — 2026-09-13 10,530 verdicts: 30 models × 93 devices (36 GPUs + 57 Mac configs) × per-platform quant tiers. Each row is generated by fitllm-engine from architecture inputs pinned to official config.json files. Runtime and OS reserves remain documented estimates. Reproduce it yourself: npm run census. Assumptions: context = min(8K, model max) · KV cache F16 · platform reserve/headroom per engine. Interactive per-combo pages: fitllm.run/can-i-run.… See the full description on the dataset page: https://huggingface.co/datasets/click6067/fitllm-fit-census.tabular10K<n<100K0 likes69 downloads18d agoHugging Face17alexandrayakovleva /single-click_bench Single-Click Benchmark for Web Interaction This benchmark defines a minimal web interaction task that requires only a single click to complete. Each instance includes two task formulations (simplified and human-like), a pre-saved HTML file for obtaining screenshots or metadata, and target annotations with the element’s bounding box and XPath for evaluation. The dataset enables systematic evaluation of web agents’ capabilities such as visual grounding, task understanding, and action… See the full description on the dataset page: https://huggingface.co/datasets/alexandrayakovleva/single-click_bench.tabularn<1K0 likes62 downloads11mo agoHugging Face18clirim911 /atlaspi-historical-geography AtlasPI — Historical Geography Dataset 1,006 historical geopolitical entities · 643 events · 55 periods · 104 dynasty chains · 252 cities · 41 trade routes The first open dataset specifically designed for AI agents working on historical geography questions. Apache 2.0 licensed. Includes real GeoJSON boundaries from academic sources, not placeholder polygons. Temporal range: 4500 BCE → 2024 CE Geographic coverage: all inhabited continents (Asia 31%, Africa 18%, Americas 17%… See the full description on the dataset page: https://huggingface.co/datasets/clirim911/atlaspi-historical-geography.tabularquestion-answering1K<n<10K0 likes50 downloads3mo agoHugging Face19agentlans /ClimbMix-sample Unofficial NVIDIA Nemotron-ClimbMix (Subsampled) This dataset is a curated, subsampled version of OptimalScale/ClimbMix, which itself is a detokenized version of NVIDIA's official pretraining dataset, nvidia/Nemotron-ClimbMix. It is designed for researchers and developers looking for a smaller, well-shuffled slice of the Nemotron pretraining data for quick experimentation, testing, or ablation studies. Processing Method To create this streamlined version, the… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/ClimbMix-sample.tabulartext-generation1M<n<10M0 likes44 downloads4mo agoHugging Face20mteb /climate-fever-v2 ClimateFEVER.v2 An MTEB dataset Massive Text Embedding Benchmark CLIMATE-FEVER is a dataset following the FEVER methodology, containing 1,535 real-world climate change claims. This updated version addresses corpus mismatches and qrel inconsistencies in MTEB, restoring labels while refining corpus-query alignment for better accuracy. Task category t2t Domains Academic, Written Reference https://www.sustainablefinance.uzh.ch/en/research/climate-fever.html… See the full description on the dataset page: https://huggingface.co/datasets/mteb/climate-fever-v2.tabulartext-retrieval10K<n<100K0 likes38 downloads1y agoHugging Face21thientrangngv /SERA-KimiK3-Django-SWEAgent-Cliff32k-T1 SERA Kimi-K3 Django SWE-Agent — Cliff-chunked T1 (first rollout) 572 training records built from 210 Kimi-K3 SWE-agent trajectories on Django, split to fit a 32,768-token context with CliffCompaction instead of being truncated. Why chunked A 100+ step agent rollout does not fit a 32k training window — 76% of the source T1 trajectories exceed it. Truncating them throws away most of the supervision, and trains the model on a context format it never sees at… See the full description on the dataset page: https://huggingface.co/datasets/thientrangngv/SERA-KimiK3-Django-SWEAgent-Cliff32k-T1.tabulartext-generationn<1K1 likes38 downloads2mo agoHugging Face22DaviBonetto /spectralbio-clinvar SpectralBio Dataset Card Public Hierarchy SpectralBio separates manuscript-facing scientific centrality from frozen executable replay centrality. Flagship scientific result: BRCA2 covariance-aware augmentation against a stronger five-model ESM-1v baseline Validation anchor: TP53 is the only frozen public canonical replay surface Breadth surface: support-ranked top-25 feasible panel derived from the 15,752-gene ClinVar scan Boundary surfaces: protocol sweep and BRCA1… See the full description on the dataset page: https://huggingface.co/datasets/DaviBonetto/spectralbio-clinvar.tabulartext-classification1K<n<10K0 likes36 downloads6mo agoHugging Face23CausalNLP /stride-preproc-climbmix STRIDE: Preprocessed ClimbMix Tokenized ClimbMix sequences used for STRIDE's pretraining-data attribution experiments. There is one training pool per released nanochat depth (d12, d16, d20, d24) and one held-out test set shared across depths. These are the corpora the released nanochat operators attribute over, so STRIDE's column indices line up with the lines of these files. Files File Sequences Size Contents climbmix_train_d12.jsonl 1,317,003 3.8 GB training… See the full description on the dataset page: https://huggingface.co/datasets/CausalNLP/stride-preproc-climbmix.tabulartext-generation10M<n<100M0 likes35 downloads3mo agoHugging Face24thientrangngv /SERA-KimiK3-Django-SWEAgent-Cliff32k-T2 SERA Kimi-K3 Django SWE-Agent — Cliff-chunked T2 (second rollout) 227 training records built from 137 Kimi-K3 SWE-agent trajectories on Django, split to fit a 32,768-token context with CliffCompaction instead of being truncated. Why chunked A 100+ step agent rollout does not fit a 32k training window — 27% of the source T2 trajectories exceed it. Truncating them throws away most of the supervision, and trains the model on a context format it never sees at… See the full description on the dataset page: https://huggingface.co/datasets/thientrangngv/SERA-KimiK3-Django-SWEAgent-Cliff32k-T2.tabulartext-generationn<1K1 likes33 downloads2mo agoHugging Face25pnu-clink /finject FInject Dataset Card FInject is a financial unanswerability benchmark built by transforming answerable financial reasoning problems into controlled unanswerable variants. Each row preserves the original question and pairs an answerable original context with a perturbed context that is no longer sufficient to support a unique answer. Dataset Summary Seed source: 78 answerable hard problems from FinanceReasoning. Final release size: 426 unanswerable variants.… See the full description on the dataset page: https://huggingface.co/datasets/pnu-clink/finject.tabularquestion-answeringn<1K0 likes31 downloads4mo agoHugging Face26zcaoyao /improved_aesthetics_6.5plus_clip_retrievaltabularn<1K0 likes30 downloads2y agoHugging Face27aliaagheis /clickhouse-server-imagetabularn<1K0 likes28 downloads2mo agoHugging Face28tanhaosheng /surgeon-tested-clinical-ai-benchmark Surgeon-Tested Clinical AI Benchmark (TH-CAB v1.1) An independent, reproducible evaluation of large language models (LLMs) on real, de-identified cancer cases — scored by a practicing surgeon item-by-item against current clinical guidelines. Homepage & full leaderboard: https://tanhaosheng.asia Methodology (citable authority, TH-CAB v1.1): https://tanhaosheng.asia/methodology/ Open data layer: https://tanhaosheng.asia/data/ This is a benchmark / research dataset, not clinical… See the full description on the dataset page: https://huggingface.co/datasets/tanhaosheng/surgeon-tested-clinical-ai-benchmark.tabulartable-to-textn<1K1 likes28 downloads2mo agoHugging Face29vickyvamsi22 /Product_reviews_click_stream_Dataatabular10K<n<100K0 likes23 downloads5mo agoHugging Face30Chtholly17 /ClinSeekAgent_RL ClinSeekAgent_RL RL-ready prompt sets derived from Letian2003/DeepMed_trajectory. The source release ships agent rollout trajectories: every question appears 4–5 times, once per sampled run, each carrying the full messages transcript and that run's outcome. That shape is built for SFT and trajectory analysis. This derivative strips it back to what an RL loop actually needs — the prompt and its ground-truth label, once per question — so the policy generates its own trajectories… See the full description on the dataset page: https://huggingface.co/datasets/Chtholly17/ClinSeekAgent_RL.tabular10K<n<100K0 likes22 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.