Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jliang097 /FARM_training_test FARM Aerial Radio Map (ARM) Dataset Paper: FARM: Foundational Aerial Radio Map for Intelligent Low-Altitude Networking (https://arxiv.org/abs/2604.17362) Overview This repository releases the constructed ARM datasets based on ARM-Omni for FARM training, in-domain evaluation (D1-D10), and zero-shot evaluation (P1, F1, and A1). The dataset coverage is summarized below: Dataset Frequencies (GHz) Max Rx Height (m) Beamwidths Map Grid Size Volume D1 2.1… See the full description on the dataset page: https://huggingface.co/datasets/jliang097/FARM_training_test.tabularimage-to-image10K<n<100K1 likes329 downloads5mo agoHugging Face02sebastian-hofstaetter /tripclick-training TripClick Baselines with Improved Training Data Establishing Strong Baselines for TripClick Health Retrieval Sebastian Hofstätter, Sophia Althammer, Mete Sertkan and Allan Hanbury https://arxiv.org/abs/2201.00365 tl;dr We create strong re-ranking and dense retrieval baselines (BERTCAT, BERTDOT, ColBERT, and TK) for TripClick (health ad-hoc retrieval). We improve the – originally too noisy – training data with a simple negative sampling policy. We achieve large gains over BM25 in the… See the full description on the dataset page: https://huggingface.co/datasets/sebastian-hofstaetter/tripclick-training.tabulartext-retrieval1M<n<10M1 likes125 downloads4y agoHugging Face03rtk-training /bjj-kimura-lesson001-trial BJJ Kimura from Side Control — Research Trial (sampled across the action arc) Tier: Research / Evaluation (free, CC BY-NC-SA 4.0) Source: RTK Motion Intelligence Platform · api.rtkmotion.io Full commercial dataset: rtk-training/bjj-kimura-lesson001 (gated) A temporally-sampled trial subset of a full 4D motion-capture session of a Brazilian Jiu-Jitsu Kimura submission from side control, demonstrated by a Former IBJJF World Champion (anonymized) with a training partner.… See the full description on the dataset page: https://huggingface.co/datasets/rtk-training/bjj-kimura-lesson001-trial.tabularrobotics10K<n<100K0 likes87 downloads2mo agoHugging Face04aitrainer-work /ai-training-labor-market AI Training Labor Market Data Pay, listing volume, listing lifespan and hiring outcomes for AI training work: paid tasks in which people write, rate and correct the output of large language models, usually as contractors hired through online platforms. Built from 9,475 job listings observed on 21 hiring platforms between December 30, 2025 and September 29, 2026. Version 2026-09, generated September 29, 2026. Online edition, updated with every site build: AI training job market… See the full description on the dataset page: https://huggingface.co/datasets/aitrainer-work/ai-training-labor-market.tabularn<1K0 likes77 downloads11d agoHugging Face05balakrishnatrt1 /ga4-dataset-from-BQ-trainingset-2monthstabular100K<n<1M0 likes70 downloads1y agoHugging Face06bhumika-tewari-282006 /ayurvedic-affinity-training-data Ayurvedic Affinity Training Data The real, processed protein-ligand binding-affinity training set used to train the companion model bhumika-tewari-282006/ayurvedic-drug-discovery-affinity-model (RandomForest regressor for pKd prediction). What this is 181 real protein-ligand complexes from PDBBind v2013-core, each with: a real PDB structure identifier (pdb_code) the real ligand SMILES the real experimentally-measured binding affinity (pkd) 39 real, computed… See the full description on the dataset page: https://huggingface.co/datasets/bhumika-tewari-282006/ayurvedic-affinity-training-data.tabulartabular-regressionn<1K0 likes70 downloads8d agoHugging Face07violetxi /chess_puzzle_training_datasets_lt-2400 Chess puzzle training datasets: rating below 2400 This is a filtered derivative of pavelslab-nyu/chess_puzzle_training_datasets. Every retained row satisfies the exact condition: Rating < 2400 Rating is the Lichess puzzle rating, not the Elo of either player in the source game. The original column names, column order, directory layout, and CSV schemas are preserved. As in the upstream repository, Hugging Face discovers all three CSVs as one default configuration with one train… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/chess_puzzle_training_datasets_lt-2400.tabulartext-generation100K<n<1M0 likes53 downloads2mo agoHugging Face08ICRA-Competitions /isaac-gr00t-ikea-training-report Isaac GR00T IKEA training report Portable export of the W&B run unitree_g1_ikea_batch32_20260730. The run stopped after a clean host shutdown; the last W&B metric is step 31,750 and the last complete checkpoint is 31,000. Summary First logged training loss: 1.4667 Last logged training loss: 0.1256 Lowest 1,000-step rolling loss: 0.1249 at step 31,750 Mean GPU compute utilization: 51.8% Mean allocated GPU memory: 29.0% Provisional checkpoint choice: 30,000… See the full description on the dataset page: https://huggingface.co/datasets/ICRA-Competitions/isaac-gr00t-ikea-training-report.document1K<n<10K1 likes50 downloads2mo agoHugging Face09electricsheepafrica /africa-synth-education-technical-vocational-training-comoros Africa Synth Education Technical Vocational Training Comoros | Africa (Electric Sheep Africa metadata inventory) Size category: 10K<n<100K - Formats: csv - Sector: economics_finance - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-education-technical-vocational-training-comoros.tabulartabular-classification10K<n<100K0 likes48 downloads2mo agoHugging Face10ehyo /GPU-Resources-Estimation-for-Deep-Learning-Training-Tasks GPUMemNet and GPUUtilNet Dataset This dataset accompanies the paper “GPU Memory and Utilization Estimation for Training-Aware Resource Management: Opportunities and Limitations.” It contains synthetic deep learning training configurations and their measured GPU memory consumption and utilization characteristics. Dataset configurations The dataset is divided into separate configurations because MLP, CNN, and Transformer workloads use different feature schemas.… See the full description on the dataset page: https://huggingface.co/datasets/ehyo/GPU-Resources-Estimation-for-Deep-Learning-Training-Tasks.tabulartabular-regression10K<n<100K0 likes47 downloads4mo agoHugging Face11aitrainer-work /ai-training-companies AI Training Companies Dataset 54 companies that pay people to train and evaluate AI models: expert networks, crowdwork platforms and data vendors. One row per company, covering how it hires (AI interviews, assessments, ID checks), how it pays (payout frequency, payment methods, eligible countries), what work it offers (RLHF, model evaluation, red teaming, coding, annotation, speech data) and what its job listings pay. Version 2026.09.1, released 2026-09-29. Online edition with… See the full description on the dataset page: https://huggingface.co/datasets/aitrainer-work/ai-training-companies.tabularn<1K0 likes47 downloads11d agoHugging Face12Geoweaver /ozone_training_data Ozone Training Data Dataset Summary The Ozone training dataset contains information about ozone levels, temperature, wind speed, pressure, and other related atmospheric variables across various geographic locations and time periods. It includes detailed daily observations from multiple data sources for comprehensive environmental and air quality analysis. Geographic coordinates (latitude and longitude) and timestamps (month, day, and hour) provide spatial and temporal… See the full description on the dataset page: https://huggingface.co/datasets/Geoweaver/ozone_training_data.tabular1M<n<10M2 likes41 downloads2y agoHugging Face13Anna4242 /grpo-training-plotstabular1K<n<10K0 likes39 downloads11mo agoHugging Face14rud11fadte /cs-trainingtabular100K<n<1M0 likes34 downloads4mo agoHugging Face15ClarusC64 /clinical-quad-investigator-turnover-training-reset-protocol-deviations-data-lag-v0.1Clinical Quad Investigator Turnover Training Reset Protocol Deviations Data Lag v0.1 Each row is a site monthly snapshot. Core quad Investigator turnoverTraining resetProtocol deviationsData lag Target label_primary_fail_next_90d Files data/train.csvdata/tester.csvscorer.py Evaluation Run model on data/tester.csvReturn predictions row alignedScore with scorer.py License MIT This dataset identifies a measurable coupling pattern associated with systemic instability. The sample demonstrates the… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-investigator-turnover-training-reset-protocol-deviations-data-lag-v0.1.tabulartext-classificationn<1K0 likes31 downloads8mo agoHugging Face16jamesdborin /Nemotron-3-Nano-RL-Training-Blend-prompt-only Nemotron-3-Nano-RL-Training-Blend-prompt-only Prompt-only extraction from nvidia/Nemotron-3-Nano-RL-Training-Blend. Files: prompts.csv: one prompt extraction record per source row. Records include prompt, separated system_prompt, and structured tools when the source row defines available tools. Nested values are JSON-encoded inside CSV cells. summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts. null_or_empty_rows.md: row indexes where… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-3-Nano-RL-Training-Blend-prompt-only.tabular10K<n<100K0 likes31 downloads3mo agoHugging Face17jamesdborin /Nemotron-RL-Ultra-Training-Blends-prompt-only Nemotron-RL-Ultra-Training-Blends-prompt-only Prompt-only extraction from nvidia/Nemotron-RL-Ultra-Training-Blends. Files: prompts.csv: one prompt extraction record per source row. Records include prompt, separated system_prompt, and structured tools when the source row defines available tools. Nested values are JSON-encoded inside CSV cells. summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts. null_or_empty_rows.md: row indexes where… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Ultra-Training-Blends-prompt-only.tabular100K<n<1M0 likes30 downloads3mo agoHugging Face18patchmedia-org /remote-ai-evaluation-training-market-snapshot Dataset Description This is an aggregate August 22, 2026 research snapshot from Specialist AI Work, an independent PatchMedia tracker of reviewed remote AI evaluation, AI training, data annotation-adjacent, and expert-review opportunities. The live Specialist AI Work inventory has advanced since this snapshot. The counts in this repository describe the immutable August 22 research object; they are not a claim about today's inventory. Reporting date: 2026-08-22 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/patchmedia-org/remote-ai-evaluation-training-market-snapshot.tabularn<1K0 likes27 downloads2mo agoHugging Face19eromang /cyberscale-training-cves CyberScale Training CVEs Training dataset for the CyberScale vulnerability severity scorer. Contains 30,641 CVEs with CVSS v3.x scores, descriptions, and CWE classifications. Schema Column Type Description cve_id string CVE identifier (e.g., CVE-2024-1234) description string Vulnerability description (English) cvss_score float CVSS v3.x base score (0.0-10.0) cvss_version string CVSS version (3.0 or 3.1) cwe string CWE identifier (e.g., CWE-79), may be… See the full description on the dataset page: https://huggingface.co/datasets/eromang/cyberscale-training-cves.tabulartext-classification10K<n<100K0 likes26 downloads7mo agoHugging Face20as-cle-bert /VirBiCla-training Dataset Card for VirBiCla-training VirBiCla is a ML-based viral DNA detector designed for long-read sequencing metagenomics. This dataset is a support dataset for training the base ML model. Dataset Details Dataset Sources [optional] Repository: GitHub repository for VirBiCla Uses This dataset is intended as support for training the base VirBiCla model Dataset Structure Dataset is a CSV file composed of 60.003 record sequences (coming… See the full description on the dataset page: https://huggingface.co/datasets/as-cle-bert/VirBiCla-training.tabular10K<n<100K1 likes25 downloads3y agoHugging Face21aitrainingjobs /ai-training-jobs-data AI Training Jobs and Pay How many AI training jobs are open, on which platforms, for which kinds of work, and what they pay per hour. AI training work is the paid work of training, evaluating and labeling data for AI models: AI trainers, RLHF evaluators, data annotators, and domain experts (law, medicine, finance, coding, math, languages) reviewing model output. The figures are measured by AITraining.jobs from the public job feeds of the platforms that hire for this work… See the full description on the dataset page: https://huggingface.co/datasets/aitrainingjobs/ai-training-jobs-data.tabularn<1K0 likes25 downloads1d agoHugging Face22GSaha567 /seq_level_training_datatabularn<1K0 likes22 downloads8mo agoHugging Face23jamesdborin /Nemotron-RL-Super-Training-Blends-prompt-only Nemotron-RL-Super-Training-Blends-prompt-only Prompt-only extraction from nvidia/Nemotron-RL-Super-Training-Blends. Files: prompts.csv: one prompt extraction record per source row. Records include prompt, separated system_prompt, and structured tools when the source row defines available tools. Nested values are JSON-encoded inside CSV cells. summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts. null_or_empty_rows.md: row indexes where… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Super-Training-Blends-prompt-only.tabular100K<n<1M0 likes18 downloads3mo agoHugging Face24jason1966 /ahmedmohamed2003_restaurant-sales-dirty-data-for-cleaning-training Restaurant Sales-Dirty Data for Cleaning Training Welcome to All Scientist Restaurant Dataset Info Source: Kaggle Original Size: 0.23 MB Kaggle Downloads: 5,163 Files: 1 Files restaurant_sales_data.csv Mirrored from Kaggle tabular10K<n<100K0 likes17 downloads6mo agoHugging Face25wlaminack /trainingnonlinear2wlaminack/nonlineartesting2 def basic(array1): x=(array1[0]-.5) y=(array1[1]-.5) z=(array1[2]-.5) t=(array1[3]-.5) r2=xx+yy+zz+tt return 7np.sin(9r2)+np.random.random()*(array1[4]-.5) f=np.apply_along_axis(basic, 1, a) tabular1M<n<10M0 likes16 downloads3y agoHugging Face26Lidor-Mashiach /mnli-training-paraphrase-augmentation MNLI Training Paraphrase Augmentation Purpose This dataset contains new paraphrases created for training paraphrase augmentation during Phase B of the research project. It was not used as an NLI evaluation set. It was not used for paraphrase consistency evaluation. It is separate from the published MNLI Paraphrase Bank used for evaluation. The generation records report zero collisions with that evaluation bank. The CSV contains only the new augmentation rows. It… See the full description on the dataset page: https://huggingface.co/datasets/Lidor-Mashiach/mnli-training-paraphrase-augmentation.tabular1M<n<10M2 likes16 downloads2mo agoHugging Face27Disclosures-SSRC /Detecting-Access-Violations-in-a-LLMs-Pre-Training-Data Beyond Public Access in LLM Pre-Training Data The official HuggingFace repository for the paper "Beyond Public Access in LLM Pre-Training Data" by The AI Disclosures Project. Using a legally obtained dataset of 34 copyrighted O'Reilly Media books, we apply the DE-COP membership inference attack method to investigate whether OpenAI's large language models were trained on copyrighted content without consent. tabular100K<n<1M0 likes14 downloads11mo agoHugging Face28Pulast /sentry_training_data Sentinel QoS Training Dataset This dataset contains synthetic network traffic features used to train the Sentry LightGBM classifier in the Sentinel-QoS project. Files training_data.csv — Tabular CSV with per-flow/session features and a target label. Columns (example) src_ip, dst_ip, src_port, dst_port protocol — e.g., TCP/UDP bytes, packets, duration app_label — human-readable application class (e.g., Video, Gaming, Browsing) target — numeric label used for model training Usage… See the full description on the dataset page: https://huggingface.co/datasets/Pulast/sentry_training_data.tabular1K<n<10K0 likes12 downloads1y agoHugging Face29eromang /cyberscale-contextual-training CyberScale Contextual Severity Training Data Training dataset for the CyberScale contextual severity classifier (Phase 2). Contains 32,000 scenarios combining CVE descriptions with NIS2 sector deployment contexts and cross-border exposure. Schema Column Type Description input_text string Formatted input: <description> [SEP] sector: <id> cross_border: <bool> score: <float> label int Severity class (0-3) sector string NIS2 sector identifier cross_border… See the full description on the dataset page: https://huggingface.co/datasets/eromang/cyberscale-contextual-training.tabulartext-classification10K<n<100K0 likes12 downloads6mo agoHugging Face30ClarusC64 /clinical-quad-device-change-measurement-drift-training-variance-endpoint-noise-v0.1Clinical Quad Device Change Measurement Drift Training Variance Endpoint Noise v0.1 Each row is a site monthly snapshot. Core quad Device changeMeasurement driftTraining varianceEndpoint noise Target label_primary_fail_next_90d Files data/train.csvdata/tester.csvscorer.py Evaluation Run model on data/tester.csvReturn predictions row alignedScore with scorer.py License MIT This dataset identifies a measurable coupling pattern associated with systemic instability. The sample demonstrates the… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-device-change-measurement-drift-training-variance-endpoint-noise-v0.1.tabulartext-classificationn<1K0 likes11 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.