datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FARM_training_test
FARM Aerial Radio Map (ARM) Dataset
Paper:
FARM: Foundational Aerial Radio Map for Intelligent Low-Altitude Networking (https://arxiv.org/abs/2604.17362)
Overview
This repository releases the constructed ARM datasets based on ARM-Omni for FARM training, in-domain evaluation (D1-D10), and zero-shot evaluation (P1, F1, and A1). The dataset coverage is summarized below:
Dataset
Frequencies (GHz)
Max Rx Height (m)
Beamwidths
Map Grid Size
Volume
D1
2.1… See the full description on the dataset page: https://huggingface.co/datasets/jliang097/FARM_training_test.tripclick-training
TripClick Baselines with Improved Training Data
Establishing Strong Baselines for TripClick Health Retrieval Sebastian Hofstätter, Sophia Althammer, Mete Sertkan and Allan Hanbury
https://arxiv.org/abs/2201.00365
tl;dr We create strong re-ranking and dense retrieval baselines (BERTCAT, BERTDOT, ColBERT, and TK) for TripClick (health ad-hoc retrieval). We improve the – originally too noisy – training data with a simple negative sampling policy. We achieve large gains over BM25 in the… See the full description on the dataset page: https://huggingface.co/datasets/sebastian-hofstaetter/tripclick-training.bjj-kimura-lesson001-trial
BJJ Kimura from Side Control — Research Trial (sampled across the action arc)
Tier: Research / Evaluation (free, CC BY-NC-SA 4.0)
Source: RTK Motion Intelligence Platform · api.rtkmotion.io
Full commercial dataset: rtk-training/bjj-kimura-lesson001 (gated)
A temporally-sampled trial subset of a full 4D motion-capture session
of a Brazilian Jiu-Jitsu Kimura submission from side control, demonstrated
by a Former IBJJF World Champion (anonymized) with a training partner.… See the full description on the dataset page: https://huggingface.co/datasets/rtk-training/bjj-kimura-lesson001-trial.ai-training-labor-market
AI Training Labor Market Data
Pay, listing volume, listing lifespan and hiring outcomes for AI training work: paid tasks in which people write, rate and correct the output of large language models, usually as contractors hired through online platforms. Built from 9,475 job listings observed on 21 hiring platforms between December 30, 2025 and September 29, 2026.
Version 2026-09, generated September 29, 2026. Online edition, updated with every site build: AI training job market… See the full description on the dataset page: https://huggingface.co/datasets/aitrainer-work/ai-training-labor-market.ga4-dataset-from-BQ-trainingset-2monthsayurvedic-affinity-training-data
Ayurvedic Affinity Training Data
The real, processed protein-ligand binding-affinity training set used to train the
companion model
bhumika-tewari-282006/ayurvedic-drug-discovery-affinity-model
(RandomForest regressor for pKd prediction).
What this is
181 real protein-ligand complexes from PDBBind v2013-core, each with:
a real PDB structure identifier (pdb_code)
the real ligand SMILES
the real experimentally-measured binding affinity (pkd)
39 real, computed… See the full description on the dataset page: https://huggingface.co/datasets/bhumika-tewari-282006/ayurvedic-affinity-training-data.chess_puzzle_training_datasets_lt-2400
Chess puzzle training datasets: rating below 2400
This is a filtered derivative of
pavelslab-nyu/chess_puzzle_training_datasets.
Every retained row satisfies the exact condition:
Rating < 2400
Rating is the Lichess puzzle rating, not the Elo of either player in the
source game. The original column names, column order, directory layout, and CSV
schemas are preserved. As in the upstream repository, Hugging Face discovers
all three CSVs as one default configuration with one train… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/chess_puzzle_training_datasets_lt-2400.isaac-gr00t-ikea-training-report
Isaac GR00T IKEA training report
Portable export of the W&B run unitree_g1_ikea_batch32_20260730. The run stopped after a clean host
shutdown; the last W&B metric is step 31,750 and the
last complete checkpoint is 31,000.
Summary
First logged training loss: 1.4667
Last logged training loss: 0.1256
Lowest 1,000-step rolling loss: 0.1249 at step 31,750
Mean GPU compute utilization: 51.8%
Mean allocated GPU memory: 29.0%
Provisional checkpoint choice: 30,000… See the full description on the dataset page: https://huggingface.co/datasets/ICRA-Competitions/isaac-gr00t-ikea-training-report.africa-synth-education-technical-vocational-training-comoros
Africa Synth Education Technical Vocational Training Comoros | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: economics_finance - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-education-technical-vocational-training-comoros.GPU-Resources-Estimation-for-Deep-Learning-Training-Tasks
GPUMemNet and GPUUtilNet Dataset
This dataset accompanies the paper
“GPU Memory and Utilization Estimation for Training-Aware Resource
Management: Opportunities and Limitations.”
It contains synthetic deep learning training configurations and their measured
GPU memory consumption and utilization characteristics.
Dataset configurations
The dataset is divided into separate configurations because MLP, CNN, and
Transformer workloads use different feature schemas.… See the full description on the dataset page: https://huggingface.co/datasets/ehyo/GPU-Resources-Estimation-for-Deep-Learning-Training-Tasks.ai-training-companies
AI Training Companies Dataset
54 companies that pay people to train and evaluate AI models: expert networks, crowdwork platforms and data vendors. One row per company, covering how it hires (AI interviews, assessments, ID checks), how it pays (payout frequency, payment methods, eligible countries), what work it offers (RLHF, model evaluation, red teaming, coding, annotation, speech data) and what its job listings pay.
Version 2026.09.1, released 2026-09-29. Online edition with… See the full description on the dataset page: https://huggingface.co/datasets/aitrainer-work/ai-training-companies.ozone_training_data
Ozone Training Data
Dataset Summary
The Ozone training dataset contains information about ozone levels, temperature, wind speed, pressure, and other related atmospheric variables across various geographic locations and time periods. It includes detailed daily observations from multiple data sources for comprehensive environmental and air quality analysis. Geographic coordinates (latitude and longitude) and timestamps (month, day, and hour) provide spatial and temporal… See the full description on the dataset page: https://huggingface.co/datasets/Geoweaver/ozone_training_data.grpo-training-plotscs-trainingclinical-quad-investigator-turnover-training-reset-protocol-deviations-data-lag-v0.1Clinical Quad Investigator Turnover Training Reset Protocol Deviations Data Lag v0.1
Each row is a site monthly snapshot.
Core quad
Investigator turnoverTraining resetProtocol deviationsData lag
Target
label_primary_fail_next_90d
Files
data/train.csvdata/tester.csvscorer.py
Evaluation
Run model on data/tester.csvReturn predictions row alignedScore with scorer.py
License
MIT
This dataset identifies a measurable coupling pattern associated with systemic instability.
The sample demonstrates the… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-investigator-turnover-training-reset-protocol-deviations-data-lag-v0.1.Nemotron-3-Nano-RL-Training-Blend-prompt-only
Nemotron-3-Nano-RL-Training-Blend-prompt-only
Prompt-only extraction from nvidia/Nemotron-3-Nano-RL-Training-Blend.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.
null_or_empty_rows.md: row indexes where… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-3-Nano-RL-Training-Blend-prompt-only.Nemotron-RL-Ultra-Training-Blends-prompt-only
Nemotron-RL-Ultra-Training-Blends-prompt-only
Prompt-only extraction from nvidia/Nemotron-RL-Ultra-Training-Blends.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.
null_or_empty_rows.md: row indexes where… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Ultra-Training-Blends-prompt-only.remote-ai-evaluation-training-market-snapshot
Dataset Description
This is an aggregate August 22, 2026 research snapshot from Specialist AI Work, an independent PatchMedia tracker of reviewed remote AI evaluation, AI training, data annotation-adjacent, and expert-review opportunities.
The live Specialist AI Work inventory has advanced since this snapshot. The counts in this repository describe the immutable August 22 research object; they are not a claim about today's inventory.
Reporting date: 2026-08-22
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/patchmedia-org/remote-ai-evaluation-training-market-snapshot.cyberscale-training-cves
CyberScale Training CVEs
Training dataset for the CyberScale vulnerability severity scorer. Contains 30,641 CVEs with CVSS v3.x scores, descriptions, and CWE classifications.
Schema
Column
Type
Description
cve_id
string
CVE identifier (e.g., CVE-2024-1234)
description
string
Vulnerability description (English)
cvss_score
float
CVSS v3.x base score (0.0-10.0)
cvss_version
string
CVSS version (3.0 or 3.1)
cwe
string
CWE identifier (e.g., CWE-79), may be… See the full description on the dataset page: https://huggingface.co/datasets/eromang/cyberscale-training-cves.VirBiCla-training
Dataset Card for VirBiCla-training
VirBiCla is a ML-based viral DNA detector designed for long-read sequencing metagenomics.
This dataset is a support dataset for training the base ML model.
Dataset Details
Dataset Sources [optional]
Repository: GitHub repository for VirBiCla
Uses
This dataset is intended as support for training the base VirBiCla model
Dataset Structure
Dataset is a CSV file composed of 60.003 record sequences (coming… See the full description on the dataset page: https://huggingface.co/datasets/as-cle-bert/VirBiCla-training.ai-training-jobs-data
AI Training Jobs and Pay
How many AI training jobs are open, on which platforms, for which kinds of work, and what they pay per hour. AI training work is the paid work of training, evaluating and labeling data for AI models: AI trainers, RLHF evaluators, data annotators, and domain experts (law, medicine, finance, coding, math, languages) reviewing model output.
The figures are measured by AITraining.jobs from the public job feeds of the platforms that hire for this work… See the full description on the dataset page: https://huggingface.co/datasets/aitrainingjobs/ai-training-jobs-data.seq_level_training_dataNemotron-RL-Super-Training-Blends-prompt-only
Nemotron-RL-Super-Training-Blends-prompt-only
Prompt-only extraction from nvidia/Nemotron-RL-Super-Training-Blends.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.
null_or_empty_rows.md: row indexes where… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Super-Training-Blends-prompt-only.ahmedmohamed2003_restaurant-sales-dirty-data-for-cleaning-training
Restaurant Sales-Dirty Data for Cleaning Training
Welcome to All Scientist Restaurant
Dataset Info
Source: Kaggle
Original Size: 0.23 MB
Kaggle Downloads: 5,163
Files: 1
Files
restaurant_sales_data.csv
Mirrored from Kaggle
trainingnonlinear2wlaminack/nonlineartesting2
def basic(array1):
x=(array1[0]-.5)
y=(array1[1]-.5)
z=(array1[2]-.5)
t=(array1[3]-.5)
r2=xx+yy+zz+tt
return 7np.sin(9r2)+np.random.random()*(array1[4]-.5)
f=np.apply_along_axis(basic, 1, a)
mnli-training-paraphrase-augmentation
MNLI Training Paraphrase Augmentation
Purpose
This dataset contains new paraphrases created for training paraphrase augmentation during Phase B of the research project.
It was not used as an NLI evaluation set.
It was not used for paraphrase consistency evaluation.
It is separate from the published MNLI Paraphrase Bank used for evaluation. The generation records report zero collisions with that evaluation bank.
The CSV contains only the new augmentation rows. It… See the full description on the dataset page: https://huggingface.co/datasets/Lidor-Mashiach/mnli-training-paraphrase-augmentation.Detecting-Access-Violations-in-a-LLMs-Pre-Training-Data
Beyond Public Access in LLM Pre-Training Data
The official HuggingFace repository for the paper "Beyond Public Access in LLM Pre-Training Data" by The AI Disclosures Project.
Using a legally obtained dataset of 34 copyrighted O'Reilly Media books, we apply the DE-COP membership inference attack method to investigate whether OpenAI's large language models were trained on copyrighted content without consent.
sentry_training_data
Sentinel QoS Training Dataset
This dataset contains synthetic network traffic features used to train the Sentry LightGBM classifier in the Sentinel-QoS project.
Files
training_data.csv — Tabular CSV with per-flow/session features and a target label.
Columns (example)
src_ip, dst_ip, src_port, dst_port
protocol — e.g., TCP/UDP
bytes, packets, duration
app_label — human-readable application class (e.g., Video, Gaming, Browsing)
target — numeric label used for model training
Usage… See the full description on the dataset page: https://huggingface.co/datasets/Pulast/sentry_training_data.cyberscale-contextual-training
CyberScale Contextual Severity Training Data
Training dataset for the CyberScale contextual severity classifier (Phase 2). Contains 32,000 scenarios combining CVE descriptions with NIS2 sector deployment contexts and cross-border exposure.
Schema
Column
Type
Description
input_text
string
Formatted input: <description> [SEP] sector: <id> cross_border: <bool> score: <float>
label
int
Severity class (0-3)
sector
string
NIS2 sector identifier
cross_border… See the full description on the dataset page: https://huggingface.co/datasets/eromang/cyberscale-contextual-training.clinical-quad-device-change-measurement-drift-training-variance-endpoint-noise-v0.1Clinical Quad Device Change Measurement Drift Training Variance Endpoint Noise v0.1
Each row is a site monthly snapshot.
Core quad
Device changeMeasurement driftTraining varianceEndpoint noise
Target
label_primary_fail_next_90d
Files
data/train.csvdata/tester.csvscorer.py
Evaluation
Run model on data/tester.csvReturn predictions row alignedScore with scorer.py
License
MIT
This dataset identifies a measurable coupling pattern associated with systemic instability.
The sample demonstrates the… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-device-change-measurement-drift-training-variance-endpoint-noise-v0.1.
