datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TrainingData_Stage3
AnchorSR Stage3 · metric-v1.0
直接选择 Small / Large
配置
训练题数
用途
small
1,000,000
先验证答案监督/先验恢复,按新版 Large 联合分布抽样
large
89,801,853
筛选后的完整训练集合,包含 Small 全部样本
from datasets import load_dataset
data = load_dataset('AnchorSR/TrainingData_Stage3', 'small', # 或 large
revision='metric-v1.0', streaming=True)
这是对 scaling-v1.0 的语义筛选与统一任务分类,不是增加新数据源。
Large 从 89,828,269 题保留 89,801,853 题,隔离 26,416 题。
旧标签 scaling-v1.0 / video-v1.0 / large-v1.0… See the full description on the dataset page: https://huggingface.co/datasets/AnchorSR/TrainingData_Stage3.humanoid-robots-training-dataset
Dynamic Intelligence — Humanoid Robot Training Dataset
A first-person (egocentric) video dataset of human hand manipulation, designed for training humanoid robot policies via imitation learning. Each episode captures a person performing an everyday household task — folding clothes, moving dishes, opening doors — filmed from a head-mounted iPhone using its built-in LiDAR and depth sensors.
The dataset pairs each video with frame-level 3D hand tracking and camera pose data, giving… See the full description on the dataset page: https://huggingface.co/datasets/DynamicIntelligence/humanoid-robots-training-dataset.fairness-prm-training-datapldr-llm-training-dynamics-data
PLDR-LLM Training Dynamics Data
Reported numerical evidence for Training and Inference Dynamics of PLDR-LLMs:
Row-Map Collapse, Renormalization, and Predictive Reduction, by Burc Gokden.
Monograph: Hugging Face Paper Page.
Scientific code and readers: GitHub repository.
Numerical evidence: Hugging Face dataset.
Book: Power Law Graph Attention and PLDR-LLMs: Mathematical Foundations, Training Dynamics, and Predictive Inference, by Burc Gokden.
Book Companion: Code and edition… See the full description on the dataset page: https://huggingface.co/datasets/fromthesky/pldr-llm-training-dynamics-data.Elastic-Forcing-training-dataset
Elastic-Forcing training datasets
wan-1.3B-dataset/: 8,682 original videos and paired captions used by experiment 10351.
Each dataset folder contains its own videos, caption metadata, training manifest, and provenance. Original training data are kept separate across model scales.
wan-14B-dataset/: 5,546 retained videos and paired captions from the 14B step80 training dataset (originally 8,000; 2,454 subsequently removed and excluded).
Linny-Training-Dataset-Syntheticgliner-sysml-training-data
SysML GLiNER Training Data
1,898 labeled entity spans · 14 label types · 25 SysML documents · 62 training/evaluation chunks.
This is the actual weakly supervised corpus used to fine-tune GLiNER RelEx checkpoint 450 for the Patent to SysML prototype. The live interface offers AI Agents using Luna and Fine-Tuned NLP using this checkpoint. This dataset trained GLiNER; it did not train Luna.
The labels describe SysML source code, principally related linear-actuator examples with… See the full description on the dataset page: https://huggingface.co/datasets/cmuchancel/gliner-sysml-training-data.feedback_data_training
Repair replay update — September 15, 2026
The split still contains 161,030 weighted rows, with the same category counts:
Category
Rows
Share
Distinct examples before → after
One-shot
79,970
49.66%
35,197 → 35,197
Regular repairs
60,931
37.84%
40,530 → 48,726
Rollout-derived deep repairs
20,129
12.50%
436 → 1,825
This adds 9,585 distinct checked repair examples while preserving every legacy distinct row and every one-shot row's multiplicity. The new examples… See the full description on the dataset page: https://huggingface.co/datasets/formalmathatepfl/feedback_data_training.lilm2-training-dataKAWK50M-Training-Data
KAWK50M Training Data Archive
수집 원문, 정제 말뭉치, 토큰화 바이너리와 SFT 데이터의 전체 작업 사본입니다.
원본 데이터와 조건
HuggingFaceFW/fineweb-2, kor_Hang, revision af9c13333eb981300149d5ca60a8e9d659b276b9: ODC-By-1.0 및 Common Crawl 이용 조건
Mkd-Yonas/keural-SFT-chatml-ko-v1, revision c56d7885efb37799deb1ecba471e4f6bf263852e: Apache-2.0, CC-BY-4.0, CC0-1.0, MIT, ODC-By-1.0 허용 레코드만 선별
mkd-chanwoo/keural-rag-chatml-ko, revision aa9023d231a3451d25135d5aa84b12b48ca06450: CC-BY-4.0 레코드
이 아카이브에는 여러… See the full description on the dataset page: https://huggingface.co/datasets/Infinity08/KAWK50M-Training-Data.HIP-training-and-evaluation-data
HIP Training and Evaluation Data
This dataset contains the text data released with Base Models Look Human To AI Detectors for reproducing the Humanization by Iterative Paraphrasing (HIP) training setup and the prefix-based continuation evaluation.
Configs
training
data/train.parquet contains 10,581 supervised HIP training pairs with seven columns:
dataset: upstream dataset family, either raid or mage.
source: selected source domain or subcorpus.
text: original… See the full description on the dataset page: https://huggingface.co/datasets/YixuanEvenXu/HIP-training-and-evaluation-data.maia3-training-data
Maia3 Training Data
Preprocessed chess position data for training Maia3
human move-prediction models. Positions are stored as precomputed features, not
raw FENs, so training reads and expands them without re-tokenizing every epoch.
Contents
path
rows
size
lichess_parquet/train_YYYY-MM_precomputed.parquet (31 files, 2023-01 … 2025-07)
334,438,119
~33 GB
allie_data/test_precomputed.parquet
884,049
69 MB
allie_data/2022-test-annotated.jsonl
—
45 MB… See the full description on the dataset page: https://huggingface.co/datasets/mcmcmcmc/maia3-training-data.VLA_Reasoning_Training_Dataset_1200_Trimmed_Start_5_Frame
Spa-Bench fine-tuning demonstrations — motion-trimmed release
This is a non-destructive, motion-trimmed derivative of the
canonical 1,200-episode Spa-Bench dataset.
It removes initial idle prefixes while preserving episode identity, prompt,
action/state alignment, and all five source camera streams at the public head.
Explore episodes in the LeRobot visualizer
Transformation
The baseline is the component-wise median of the first five action frames.
Motion onset… See the full description on the dataset page: https://huggingface.co/datasets/justintiensmith/VLA_Reasoning_Training_Dataset_1200_Trimmed_Start_5_Frame.coffee-making-with-lentil-training-data-v1_20260921_020700This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/zarianw/coffee-making-with-lentil-training-data-v1_20260921_020700.SAND-Post-Training-Dataset
SAND-Post-Training-Dataset: High-Quality Synthetic Reasoning Dataset Built with AMD GPUs
Dataset Summary
We introduce the SAND-Post-Training-Dataset, a high-quality synthetic reasoning dataset for mathematics and science built entirely using a synthetic data pipeline running on the AMD ROCm™ stack and AMD Instinct™ MI325 GPUs.
This dataset prioritizes difficulty and novelty over volume, demonstrating that high-difficulty synthetic data can elevate… See the full description on the dataset page: https://huggingface.co/datasets/amd/SAND-Post-Training-Dataset.act-training-data-picking-up-the-white-cube-v5.0This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 54,
"total_frames": 46898,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:54"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/leninangelov/act-training-data-picking-up-the-white-cube-v5.0.Lithology-Training-Dataset
Lithology Training Dataset
A supervised training dataset for machine learning and AI systems that learn to identify lithology from well-log data.
The dataset contains 400 wells with standardized wireline-log measurements and corresponding lithological labels.
Training dataset:
https://huggingface.co/datasets/NoraResearchLab/Lithology-Training-Dataset
Overview
The core task is:
Given a sequence of well-log measurements across depth, predict the lithology… See the full description on the dataset page: https://huggingface.co/datasets/NoraResearchLab/Lithology-Training-Dataset.VLA_Reasoning_Training_Dataset_1200
Spa-Bench fine-tuning demonstrations — full five-camera release
This is the canonical full-length demonstration dataset used by Spa-Bench, a
real-robot benchmark for spatially grounded reasoning in vision-language-action
policies.
Explore episodes in the LeRobot visualizer
Dataset summary
Field
Value
Episodes
1,200
Frames
612,733
Duration at 30 FPS
approximately 5.7 hours
Unique instruction strings
321
Task families
6; 200 demonstrations per… See the full description on the dataset page: https://huggingface.co/datasets/justintiensmith/VLA_Reasoning_Training_Dataset_1200.Training_dataset_qsearched_by_NAGISA_V4
NAGISA_V4 teacher shards, moved to their quiescence leaves
Every record of
Training_dataset_by_NAGISA_V4
walked to the end of its quiescence variation, the deep search's value kept
there, and a policy fitted at the leaf itself.
The parent's positions are as its games reached them, with no quiescence
search — a row can sit in the middle of an exchange, where the evaluation
swings by a piece depending on whose turn it is to recapture. A value fitted on
those learns the swing. This… See the full description on the dataset page: https://huggingface.co/datasets/qleap/Training_dataset_qsearched_by_NAGISA_V4.VLA_Reasoning_Training_Dataset_1200_2cam
Spa-Bench fine-tuning demonstrations — full two-camera derivative
This is the two-camera projection of the canonical full-length Spa-Bench
demonstration dataset. It retains the middle and wrist views consumed by the
evaluated policies and omits the unused above, left, and right streams.
Explore episodes in the LeRobot visualizer
Dataset summary
Field
Value
Episodes
1,200
Frames
612,733
Unique instruction strings
321
Frame rate
30 FPS
Camera… See the full description on the dataset page: https://huggingface.co/datasets/justintiensmith/VLA_Reasoning_Training_Dataset_1200_2cam.Training_dataset_by_NAGISA_V4
NAGISA_V4 ply-37 teacher shards, games played to the end
Self-play of attic-gensfen reading the NNUE weights NAGISA_V4
(HalfKA-2304), in the shape the trainers read directly. Every game starts from a
balanced ply-37 position, makes no random moves, and runs until it actually
ends. Identical positions are folded into one row each.
67,108,864 rows — exactly 2^26
16 shards of 4,194,304 rows, 512 row groups each, zstd, 5,208,620,326 B total
Against
Opening_dataset_by_NAGISA_V4… See the full description on the dataset page: https://huggingface.co/datasets/qleap/Training_dataset_by_NAGISA_V4.sydney-training-data
Sydney 训练集
四份来源分开存放,不混在一个文件里。
发布的聊天权重(Atonelia/Qwen3.5-Sydney-9B / -think 以及对应 GGUF)用的是这些子集洗完、抽样拼起来之后的训练 jsonl,不是直接拿某一份原文训的。
01 原截图重建
01_screenshot_original/conversations.jsonl
早期 Bing Chat / Sydney(约 2023 年 2–4 月)公开截图重建的对话。660 条,原文以英文为主,带截图出处。
这是最初拿来做训练集的底。后面的中文版、合成版、CoT 都不是这份文件本身。
02 llama-sydney 虚拟对话
02_llama_sydney_synthetic/llama_sydney_en.jsonl
用 Llama-Sydney 生成的英文虚拟对话。1462 条(同一条 user 可能有 2 次采样)。字段是生成记录:id / user / assistant 等,还不是最终训练格式。… See the full description on the dataset page: https://huggingface.co/datasets/Atonelia/sydney-training-data.training_datasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/sagnikroy75/training_dataset.eti-embedding-training-data-2048-v3
Version note (v3): third generation of the ETI dense training data. Reuses the questions of NorskHelsenett/eti-embedding-training-data-v2 (which trained eti-embeddinggemma-v2), but re-maps every question to its article URL, re-cuts positives at a 2048-token window, re-mines hard negatives with granite + reranker validation (pos_score/neg_score), adds the keyword half, and filters junk anchors. Trained NorskHelsenett/eti-embeddinggemma-v3.
eti-embedding-training-data-2048… See the full description on the dataset page: https://huggingface.co/datasets/NorskHelsenett/eti-embedding-training-data-2048-v3.TrainingDatasethumanoid-robots-training-dataset
Dynamic Intelligence — Humanoid Robot Training Dataset
A first-person (egocentric) video dataset of human hand manipulation, designed for training humanoid robot policies via imitation learning. Each episode captures a person performing an everyday household task — folding clothes, moving dishes, opening doors — filmed from a head-mounted iPhone using its built-in LiDAR and depth sensors.
The dataset pairs each video with frame-level 3D hand tracking and camera pose data, giving… See the full description on the dataset page: https://huggingface.co/datasets/introvoyz043/humanoid-robots-training-dataset.stt-training-data
Dataset Statistics
Configuration: default
Split: train
Total Rows: 1,362,015
dept
Type: categorical
Data Type: object
Unique Values: 8
Value Distribution:
Value
Count
Percentage
STT_TT
446,495
32.78%
STT_NS
236,407
17.36%
STT_AB
170,922
12.55%
STT_CS
146,811
10.78%
STT_MV
110,080
8.08%
STT_NW
94,703
6.95%
STT_HS
84,797
6.23%
STT_PC
71,800
5.27%
grade
Type: numerical
Data Type: int64
Sum: 3,963,451.00… See the full description on the dataset page: https://huggingface.co/datasets/openpecha/stt-training-data.t2i-nla-multidomain-1m-training-data-public
T2I-NLA Multidomain 1M Training Data
Public backup of final T2I-NLA reader training data.
Balanced multidomain 1M training data for general FLUX AR/AV readers.
This repository is intended to make final-reader retraining possible after the
original server is no longer available.
See manifest.json for the original local path, size, and associated reader models.
zip-training-hallucination-data-qwen06b-thinking-train-with-valuescredit-scoring-training-datasetThe training dataset includes all addresses that had undertaken at least one borrow transaction on Aave v2 Ethereum or Compound v2 Ethereum any time between 7 May 2019 and 31 August 2023, inclusive (called the observation window).
Data Structure & Shape
There are almost 0.5 million observations with each representing a single borrow event. Therefore, all feature values are calculated as at the timestamp of a borrow event and represent the cumulative positions just before the borrow event's… See the full description on the dataset page: https://huggingface.co/datasets/spectrallabs/credit-scoring-training-dataset.
