Team Ai
23 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01fromthesky /pldr-llm-training-dynamics-data PLDR-LLM Training Dynamics Data Reported numerical evidence for Training and Inference Dynamics of PLDR-LLMs: Row-Map Collapse, Renormalization, and Predictive Reduction, by Burc Gokden. Monograph: Hugging Face Paper Page. Scientific code and readers: GitHub repository. Numerical evidence: Hugging Face dataset. Book: Power Law Graph Attention and PLDR-LLMs: Mathematical Foundations, Training Dynamics, and Predictive Inference, by Burc Gokden. Book Companion: Code and edition… See the full description on the dataset page: https://huggingface.co/datasets/fromthesky/pldr-llm-training-dynamics-data.tabularothern<1K0 likes868 downloads4d agoHugging Face02liuyueyi-8 /Elastic-Forcing-training-dataset Elastic-Forcing training datasets wan-1.3B-dataset/: 8,682 original videos and paired captions used by experiment 10351. Each dataset folder contains its own videos, caption metadata, training manifest, and provenance. Original training data are kept separate across model scales. wan-14B-dataset/: 5,546 retained videos and paired captions from the 14B step80 training dataset (originally 8,000; 2,454 subsequently removed and excluded). tabular10K<n<100K1 likes757 downloads12d agoHugging Face03cmuchancel /gliner-sysml-training-data SysML GLiNER Training Data 1,898 labeled entity spans · 14 label types · 25 SysML documents · 62 training/evaluation chunks. This is the actual weakly supervised corpus used to fine-tune GLiNER RelEx checkpoint 450 for the Patent to SysML prototype. The live interface offers AI Agents using Luna and Fine-Tuned NLP using this checkpoint. This dataset trained GLiNER; it did not train Luna. The labels describe SysML source code, principally related linear-actuator examples with… See the full description on the dataset page: https://huggingface.co/datasets/cmuchancel/gliner-sysml-training-data.tabulartoken-classification1K<n<10K0 likes729 downloads6d agoHugging Face04amd /SAND-Post-Training-Dataset SAND-Post-Training-Dataset: High-Quality Synthetic Reasoning Dataset Built with AMD GPUs Dataset Summary We introduce the SAND-Post-Training-Dataset, a high-quality synthetic reasoning dataset for mathematics and science built entirely using a synthetic data pipeline running on the AMD ROCm™ stack and AMD Instinct™ MI325 GPUs. This dataset prioritizes difficulty and novelty over volume, demonstrating that high-difficulty synthetic data can elevate… See the full description on the dataset page: https://huggingface.co/datasets/amd/SAND-Post-Training-Dataset.tabularquestion-answering10K<n<100K3 likes233 downloads10mo agoHugging Face05Atonelia /sydney-training-data Sydney 训练集 四份来源分开存放,不混在一个文件里。 发布的聊天权重(Atonelia/Qwen3.5-Sydney-9B / -think 以及对应 GGUF)用的是这些子集洗完、抽样拼起来之后的训练 jsonl,不是直接拿某一份原文训的。 01 原截图重建 01_screenshot_original/conversations.jsonl 早期 Bing Chat / Sydney(约 2023 年 2–4 月)公开截图重建的对话。660 条,原文以英文为主,带截图出处。 这是最初拿来做训练集的底。后面的中文版、合成版、CoT 都不是这份文件本身。 02 llama-sydney 虚拟对话 02_llama_sydney_synthetic/llama_sydney_en.jsonl 用 Llama-Sydney 生成的英文虚拟对话。1462 条(同一条 user 可能有 2 次采样)。字段是生成记录:id / user / assistant 等,还不是最终训练格式。… See the full description on the dataset page: https://huggingface.co/datasets/Atonelia/sydney-training-data.tabulartext-generation1K<n<10K1 likes137 downloads20d agoHugging Face06cs552-the-expendables /mts-rl-training-data Source coverage MTS-Dialog contains 1,701 source dialogues. Fact generation succeeded for 1,700; one training dialogue was excluded after no valid fact record could be generated. Fact records are available for all 200 dialogues in test1. This fact-generation exclusion is separate from the source-dialogue eligibility rule used by the current G-Eval comparison. PatientAgent MTS-Dialog training data Facts and preference-training data used by the current PatientAgent… See the full description on the dataset page: https://huggingface.co/datasets/cs552-the-expendables/mts-rl-training-data.tabular10K<n<100K0 likes72 downloads2mo agoHugging Face07Training-Datasmith /k3-sft-cc0-flan Dataset Card for K3 SFT CC0 FLAN 844-row Kimi K3 synthetic instruction-tuning shard built from DPI-traced CC0/public-domain FLAN prompts in the Tülu mix. Four overlapping Hub configs expose different cohort views; adaptive is the recommended default for quality-conscious SFT mixing. Dataset Details Curated by: Training Datasmith Teacher: kimi-k3 via deltafin (local inference) Languages: English prompts; translation pairs include German, Spanish, Czech, Igbo… See the full description on the dataset page: https://huggingface.co/datasets/Training-Datasmith/k3-sft-cc0-flan.tabulartext-classification1K<n<10K0 likes58 downloads21d agoHugging Face08GIZ /vulnerability_training_data_fulltabulartext-classificationn<1K0 likes46 downloads3y agoHugging Face09hozifa1 /Training-Ai-Islamic-Dataset 🕌 Training AI Islamic Dataset 18.7M passages from classical Islamic books spanning 1,400 years of scholarship. Comprehensive collection of Islamic texts covering Quran, Hadith, Fiqh, Tafsir, Aqeedah, Seerah, Usul al-Fiqh, and Arabic Language — structured with scholarly metadata for RAG and LLM training. 📊 Dataset Structure collections/: Categorized Islamic passages compressed in JSONL format. metadata/: Scholarly master catalogs, author biographical death… See the full description on the dataset page: https://huggingface.co/datasets/hozifa1/Training-Ai-Islamic-Dataset.tabularquestion-answering10M<n<100M0 likes43 downloads1mo agoHugging Face10zzhaobz /cembra-training-data Cembra OA pseudo-SNP training data Private, model-ready data for the Cembra pseudo-SNP SVD64 model. This package contains the exact reference/alternate sequence pairs consumed by the selected development model; it contains no individual-level genotype or phenotype records. Configurations supervised 168 OA variant examples: 56 positives and 112 matched controls. 56 complete matched triplets in 49 guarded independence components. Exact five-fold… See the full description on the dataset page: https://huggingface.co/datasets/zzhaobz/cembra-training-data.tabulartext-classification10K<n<100K0 likes33 downloads2mo agoHugging Face11Jordine /patina3-v3-training-data PATINA-3 v3 — training data backup (updated 2026-09-08) Inputs of the PATINA-3 experiment (design of record: patina3/SPEC.md v3 in github.com/Jordine/entanglement_engineering). Question: does a spec's explanation have to be TRUE, or only STATED, for its value to generalize? Base: Llama-3.1-8B; upstream: Model Spec Midtraining (arXiv 2605.02087). template/tmpl_afford_full_r9_1.jsonl — the 4,600 skeletons every corpus is instantiated from; template/excluded_skeletons_r9_1_v3.json… See the full description on the dataset page: https://huggingface.co/datasets/Jordine/patina3-v3-training-data.tabular10K<n<100K0 likes33 downloads1mo agoHugging Face12milkkarten /pokemon-training-datatabular10M<n<100M0 likes25 downloads9mo agoHugging Face13Rayugacodes /kernelx-training-datatabular100K<n<1M0 likes24 downloads6mo agoHugging Face14agcbench-2026 /AGC-Judge-Training-Data AGC-Judge Training Data This dataset contains the chat-format supervision and evaluation splits used to train and validate AGC-Judge, the open-weight scorer released with AGC-Bench (Artificial General Creativity Benchmark). Each row in the messages config is a three-message chat conversation: system: scoring instruction for AGC-Judge. user: benchmark rubric, benchmark prompt, and model response to score. assistant: the JRT-corrected integer score used as the gold target. The… See the full description on the dataset page: https://huggingface.co/datasets/agcbench-2026/AGC-Judge-Training-Data.tabulartext-generation100K<n<1M0 likes22 downloads5mo agoHugging Face15gsaltintas /training_data_detokenizedtabular100K<n<1M0 likes21 downloads1y agoHugging Face16anonymous-dart-2026 /training_datasettabular10K<n<100K0 likes19 downloads6mo agoHugging Face17aimosprite /training-data-oss120btabular1K<n<10K0 likes15 downloads6mo agoHugging Face18Lonelyguyse1 /halide-training-data Project Halide Training Data Film defect detection training data for MiniCPM-V 4.6 fine-tuning. Dataset FilmDamageSimulator (Eurographics 2023) 10 film scans (4K resolution) 12,137 defect annotations 5 classes: dust, dirt, short_hair, long_hair, scratch All bounding boxes normalized to [0.0-1.0] Format JSONL with structure: Classes Class Count Color dust 7,631 Red dirt 2,700 Orange short_hair 1,341 Cyan long_hair… See the full description on the dataset page: https://huggingface.co/datasets/Lonelyguyse1/halide-training-data.tabularn<1K0 likes15 downloads4mo agoHugging Face19spkc83 /retail-bank-router-training-data Retail Bank dual-head router data Governed classifier-only data derived from PolyAI Banking77 and UCI CLINC150. It is not included in generative SFT. Train rows: 44432 Validation rows: 8589 Test rows: 16260 Domain labels: OOD=0, supported retail banking=1 Supported domain includes greetings, thanks, goodbyes, and assistant-identity questions Intent labels: 77 Banking77 intents; -100 means no intent supervision Licenses: CC-BY-4.0 See manifest.json for source revisions, hashes… See the full description on the dataset page: https://huggingface.co/datasets/spkc83/retail-bank-router-training-data.tabulartext-classification10K<n<100K0 likes15 downloads2mo agoHugging Face20quintana42 /gang-of-four-training-data Gang of Four Training Data Training data for the Gang of Four neural AI, generated from ExpertStrategy self-play. Dataset Details Format: JSONL (one JSON object per line) Size: ~1M game states Source: ExpertStrategy vs ExpertStrategy games Schema Each line contains: { "state": [328 floats], "action_mask": [40 floats], "action_idx": int, "declared_last_card": bool } Usage from training.dataset import GangOfFourDataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/quintana42/gang-of-four-training-data.tabular1M<n<10M0 likes14 downloads9mo agoHugging Face21vsamuel /Ruler_Training_Datatabular10K<n<100K0 likes8 downloads2y agoHugging Face22Rubyando59 /xenc-training-datatabular1M<n<10M0 likes6 downloads8mo agoHugging Face23bwirth /spider-classifier-training-data Spider Classifier — Training Manifest Public release of the training manifest used to fine-tune the Spiders of New Hampshire species classifier. This manifest enumerates every photo used to train, validate, and test the model. Each row links back to the original observation and photo on iNaturalist, preserving full attribution and license metadata. Source Model run: 20260528_104624_licensed_dinov2_l_14_reg4_518 Generated: 2026-05-29T02:17:00.691986+00:00… See the full description on the dataset page: https://huggingface.co/datasets/bwirth/spider-classifier-training-data.tabularn<1K0 likes6 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.