Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AnchorSR /TrainingData_Stage3 AnchorSR Stage3 · metric-v1.0 直接选择 Small / Large 配置 训练题数 用途 small 1,000,000 先验证答案监督/先验恢复,按新版 Large 联合分布抽样 large 89,801,853 筛选后的完整训练集合,包含 Small 全部样本 from datasets import load_dataset data = load_dataset('AnchorSR/TrainingData_Stage3', 'small', # 或 large revision='metric-v1.0', streaming=True) 这是对 scaling-v1.0 的语义筛选与统一任务分类,不是增加新数据源。 Large 从 89,828,269 题保留 89,801,853 题,隔离 26,416 题。 旧标签 scaling-v1.0 / video-v1.0 / large-v1.0… See the full description on the dataset page: https://huggingface.co/datasets/AnchorSR/TrainingData_Stage3.tabularvisual-question-answering100M<n<1B0 likes11k downloads18d agoHugging Face02DynamicIntelligence /humanoid-robots-training-dataset Dynamic Intelligence — Humanoid Robot Training Dataset A first-person (egocentric) video dataset of human hand manipulation, designed for training humanoid robot policies via imitation learning. Each episode captures a person performing an everyday household task — folding clothes, moving dishes, opening doors — filmed from a head-mounted iPhone using its built-in LiDAR and depth sensors. The dataset pairs each video with frame-level 3D hand tracking and camera pose data, giving… See the full description on the dataset page: https://huggingface.co/datasets/DynamicIntelligence/humanoid-robots-training-dataset.tabularrobotics10K<n<100K0 likes3.1k downloads7mo agoHugging Face03zarahall /fairness-prm-training-datatabular100K<n<1M2 likes1.3k downloads1y agoHugging Face04fromthesky /pldr-llm-training-dynamics-data PLDR-LLM Training Dynamics Data Reported numerical evidence for Training and Inference Dynamics of PLDR-LLMs: Row-Map Collapse, Renormalization, and Predictive Reduction, by Burc Gokden. Monograph: Hugging Face Paper Page. Scientific code and readers: GitHub repository. Numerical evidence: Hugging Face dataset. Book: Power Law Graph Attention and PLDR-LLMs: Mathematical Foundations, Training Dynamics, and Predictive Inference, by Burc Gokden. Book Companion: Code and edition… See the full description on the dataset page: https://huggingface.co/datasets/fromthesky/pldr-llm-training-dynamics-data.tabularothern<1K0 likes868 downloads4d agoHugging Face05liuyueyi-8 /Elastic-Forcing-training-dataset Elastic-Forcing training datasets wan-1.3B-dataset/: 8,682 original videos and paired captions used by experiment 10351. Each dataset folder contains its own videos, caption metadata, training manifest, and provenance. Original training data are kept separate across model scales. wan-14B-dataset/: 5,546 retained videos and paired captions from the 14B step80 training dataset (originally 8,000; 2,454 subsequently removed and excluded). tabular10K<n<100K1 likes757 downloads12d agoHugging Face06Lucien-shark /Linny-Training-Dataset-Synthetictabularn<1K0 likes732 downloads1h agoHugging Face07cmuchancel /gliner-sysml-training-data SysML GLiNER Training Data 1,898 labeled entity spans · 14 label types · 25 SysML documents · 62 training/evaluation chunks. This is the actual weakly supervised corpus used to fine-tune GLiNER RelEx checkpoint 450 for the Patent to SysML prototype. The live interface offers AI Agents using Luna and Fine-Tuned NLP using this checkpoint. This dataset trained GLiNER; it did not train Luna. The labels describe SysML source code, principally related linear-actuator examples with… See the full description on the dataset page: https://huggingface.co/datasets/cmuchancel/gliner-sysml-training-data.tabulartoken-classification1K<n<10K0 likes729 downloads5d agoHugging Face08formalmathatepfl /feedback_data_training Repair replay update — September 15, 2026 The split still contains 161,030 weighted rows, with the same category counts: Category Rows Share Distinct examples before → after One-shot 79,970 49.66% 35,197 → 35,197 Regular repairs 60,931 37.84% 40,530 → 48,726 Rollout-derived deep repairs 20,129 12.50% 436 → 1,825 This adds 9,585 distinct checked repair examples while preserving every legacy distinct row and every one-shot row's multiplicity. The new examples… See the full description on the dataset page: https://huggingface.co/datasets/formalmathatepfl/feedback_data_training.tabular1M<n<10M1 likes507 downloads25d agoHugging Face09glouriousgautam /lilm2-training-datatabular10M<n<100M0 likes392 downloads23d agoHugging Face10Infinity08 /KAWK50M-Training-Data KAWK50M Training Data Archive 수집 원문, 정제 말뭉치, 토큰화 바이너리와 SFT 데이터의 전체 작업 사본입니다. 원본 데이터와 조건 HuggingFaceFW/fineweb-2, kor_Hang, revision af9c13333eb981300149d5ca60a8e9d659b276b9: ODC-By-1.0 및 Common Crawl 이용 조건 Mkd-Yonas/keural-SFT-chatml-ko-v1, revision c56d7885efb37799deb1ecba471e4f6bf263852e: Apache-2.0, CC-BY-4.0, CC0-1.0, MIT, ODC-By-1.0 허용 레코드만 선별 mkd-chanwoo/keural-rag-chatml-ko, revision aa9023d231a3451d25135d5aa84b12b48ca06450: CC-BY-4.0 레코드 이 아카이브에는 여러… See the full description on the dataset page: https://huggingface.co/datasets/Infinity08/KAWK50M-Training-Data.tabularn<1K0 likes350 downloads2mo agoHugging Face11YixuanEvenXu /HIP-training-and-evaluation-data HIP Training and Evaluation Data This dataset contains the text data released with Base Models Look Human To AI Detectors for reproducing the Humanization by Iterative Paraphrasing (HIP) training setup and the prefix-based continuation evaluation. Configs training data/train.parquet contains 10,581 supervised HIP training pairs with seven columns: dataset: upstream dataset family, either raid or mage. source: selected source domain or subcorpus. text: original… See the full description on the dataset page: https://huggingface.co/datasets/YixuanEvenXu/HIP-training-and-evaluation-data.tabulartext-generation10K<n<100K1 likes324 downloads5mo agoHugging Face12mcmcmcmc /maia3-training-data Maia3 Training Data Preprocessed chess position data for training Maia3 human move-prediction models. Positions are stored as precomputed features, not raw FENs, so training reads and expands them without re-tokenizing every epoch. Contents path rows size lichess_parquet/train_YYYY-MM_precomputed.parquet (31 files, 2023-01 … 2025-07) 334,438,119 ~33 GB allie_data/test_precomputed.parquet 884,049 69 MB allie_data/2022-test-annotated.jsonl — 45 MB… See the full description on the dataset page: https://huggingface.co/datasets/mcmcmcmc/maia3-training-data.tabularother100M<n<1B1 likes321 downloads15d agoHugging Face13justintiensmith /VLA_Reasoning_Training_Dataset_1200_Trimmed_Start_5_Frame Spa-Bench fine-tuning demonstrations — motion-trimmed release This is a non-destructive, motion-trimmed derivative of the canonical 1,200-episode Spa-Bench dataset. It removes initial idle prefixes while preserving episode identity, prompt, action/state alignment, and all five source camera streams at the public head. Explore episodes in the LeRobot visualizer Transformation The baseline is the component-wise median of the first five action frames. Motion onset… See the full description on the dataset page: https://huggingface.co/datasets/justintiensmith/VLA_Reasoning_Training_Dataset_1200_Trimmed_Start_5_Frame.tabularrobotics100K<n<1M0 likes250 downloads1mo agoHugging Face14zarianw /coffee-making-with-lentil-training-data-v1_20260921_020700This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/zarianw/coffee-making-with-lentil-training-data-v1_20260921_020700.tabularrobotics100K<n<1M0 likes240 downloads5d agoHugging Face15amd /SAND-Post-Training-Dataset SAND-Post-Training-Dataset: High-Quality Synthetic Reasoning Dataset Built with AMD GPUs Dataset Summary We introduce the SAND-Post-Training-Dataset, a high-quality synthetic reasoning dataset for mathematics and science built entirely using a synthetic data pipeline running on the AMD ROCm™ stack and AMD Instinct™ MI325 GPUs. This dataset prioritizes difficulty and novelty over volume, demonstrating that high-difficulty synthetic data can elevate… See the full description on the dataset page: https://huggingface.co/datasets/amd/SAND-Post-Training-Dataset.tabularquestion-answering10K<n<100K3 likes233 downloads10mo agoHugging Face16leninangelov /act-training-data-picking-up-the-white-cube-v5.0This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so_follower", "total_episodes": 54, "total_frames": 46898, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:54" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/leninangelov/act-training-data-picking-up-the-white-cube-v5.0.tabularrobotics10K<n<100K0 likes222 downloads18d agoHugging Face17NoraResearchLab /Lithology-Training-Dataset Lithology Training Dataset A supervised training dataset for machine learning and AI systems that learn to identify lithology from well-log data. The dataset contains 400 wells with standardized wireline-log measurements and corresponding lithological labels. Training dataset: https://huggingface.co/datasets/NoraResearchLab/Lithology-Training-Dataset Overview The core task is: Given a sequence of well-log measurements across depth, predict the lithology… See the full description on the dataset page: https://huggingface.co/datasets/NoraResearchLab/Lithology-Training-Dataset.tabularother1M<n<10M1 likes159 downloads7d agoHugging Face18justintiensmith /VLA_Reasoning_Training_Dataset_1200 Spa-Bench fine-tuning demonstrations — full five-camera release This is the canonical full-length demonstration dataset used by Spa-Bench, a real-robot benchmark for spatially grounded reasoning in vision-language-action policies. Explore episodes in the LeRobot visualizer Dataset summary Field Value Episodes 1,200 Frames 612,733 Duration at 30 FPS approximately 5.7 hours Unique instruction strings 321 Task families 6; 200 demonstrations per… See the full description on the dataset page: https://huggingface.co/datasets/justintiensmith/VLA_Reasoning_Training_Dataset_1200.tabularrobotics100K<n<1M0 likes148 downloads1mo agoHugging Face19qleap /Training_dataset_qsearched_by_NAGISA_V4 NAGISA_V4 teacher shards, moved to their quiescence leaves Every record of Training_dataset_by_NAGISA_V4 walked to the end of its quiescence variation, the deep search's value kept there, and a policy fitted at the leaf itself. The parent's positions are as its games reached them, with no quiescence search — a row can sit in the middle of an exchange, where the evaluation swings by a piece depending on whose turn it is to recapture. A value fitted on those learns the swing. This… See the full description on the dataset page: https://huggingface.co/datasets/qleap/Training_dataset_qsearched_by_NAGISA_V4.tabularreinforcement-learning10M<n<100M0 likes146 downloads28d agoHugging Face20justintiensmith /VLA_Reasoning_Training_Dataset_1200_2cam Spa-Bench fine-tuning demonstrations — full two-camera derivative This is the two-camera projection of the canonical full-length Spa-Bench demonstration dataset. It retains the middle and wrist views consumed by the evaluated policies and omits the unused above, left, and right streams. Explore episodes in the LeRobot visualizer Dataset summary Field Value Episodes 1,200 Frames 612,733 Unique instruction strings 321 Frame rate 30 FPS Camera… See the full description on the dataset page: https://huggingface.co/datasets/justintiensmith/VLA_Reasoning_Training_Dataset_1200_2cam.tabularrobotics100K<n<1M0 likes140 downloads1mo agoHugging Face21qleap /Training_dataset_by_NAGISA_V4 NAGISA_V4 ply-37 teacher shards, games played to the end Self-play of attic-gensfen reading the NNUE weights NAGISA_V4 (HalfKA-2304), in the shape the trainers read directly. Every game starts from a balanced ply-37 position, makes no random moves, and runs until it actually ends. Identical positions are folded into one row each. 67,108,864 rows — exactly 2^26 16 shards of 4,194,304 rows, 512 row groups each, zstd, 5,208,620,326 B total Against Opening_dataset_by_NAGISA_V4… See the full description on the dataset page: https://huggingface.co/datasets/qleap/Training_dataset_by_NAGISA_V4.tabularreinforcement-learning10M<n<100M0 likes138 downloads1mo agoHugging Face22Atonelia /sydney-training-data Sydney 训练集 四份来源分开存放,不混在一个文件里。 发布的聊天权重(Atonelia/Qwen3.5-Sydney-9B / -think 以及对应 GGUF)用的是这些子集洗完、抽样拼起来之后的训练 jsonl,不是直接拿某一份原文训的。 01 原截图重建 01_screenshot_original/conversations.jsonl 早期 Bing Chat / Sydney(约 2023 年 2–4 月)公开截图重建的对话。660 条,原文以英文为主,带截图出处。 这是最初拿来做训练集的底。后面的中文版、合成版、CoT 都不是这份文件本身。 02 llama-sydney 虚拟对话 02_llama_sydney_synthetic/llama_sydney_en.jsonl 用 Llama-Sydney 生成的英文虚拟对话。1462 条(同一条 user 可能有 2 次采样)。字段是生成记录:id / user / assistant 等,还不是最终训练格式。… See the full description on the dataset page: https://huggingface.co/datasets/Atonelia/sydney-training-data.tabulartext-generation1K<n<10K1 likes137 downloads20d agoHugging Face23sagnikroy75 /training_datasetThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/sagnikroy75/training_dataset.tabularrobotics10K<n<100K0 likes126 downloads21d agoHugging Face24NorskHelsenett /eti-embedding-training-data-2048-v3 Version note (v3): third generation of the ETI dense training data. Reuses the questions of NorskHelsenett/eti-embedding-training-data-v2 (which trained eti-embeddinggemma-v2), but re-maps every question to its article URL, re-cuts positives at a 2048-token window, re-mines hard negatives with granite + reranker validation (pos_score/neg_score), adds the keyword half, and filters junk anchors. Trained NorskHelsenett/eti-embeddinggemma-v3. eti-embedding-training-data-2048… See the full description on the dataset page: https://huggingface.co/datasets/NorskHelsenett/eti-embedding-training-data-2048-v3.tabularsentence-similarity100K<n<1M0 likes115 downloads1mo agoHugging Face25Haesteining /TrainingDatasettabular1M<n<10M0 likes101 downloads2y agoHugging Face26introvoyz043 /humanoid-robots-training-dataset Dynamic Intelligence — Humanoid Robot Training Dataset A first-person (egocentric) video dataset of human hand manipulation, designed for training humanoid robot policies via imitation learning. Each episode captures a person performing an everyday household task — folding clothes, moving dishes, opening doors — filmed from a head-mounted iPhone using its built-in LiDAR and depth sensors. The dataset pairs each video with frame-level 3D hand tracking and camera pose data, giving… See the full description on the dataset page: https://huggingface.co/datasets/introvoyz043/humanoid-robots-training-dataset.tabularrobotics10K<n<100K0 likes101 downloads26d agoHugging Face27openpecha /stt-training-data Dataset Statistics Configuration: default Split: train Total Rows: 1,362,015 dept Type: categorical Data Type: object Unique Values: 8 Value Distribution: Value Count Percentage STT_TT 446,495 32.78% STT_NS 236,407 17.36% STT_AB 170,922 12.55% STT_CS 146,811 10.78% STT_MV 110,080 8.08% STT_NW 94,703 6.95% STT_HS 84,797 6.23% STT_PC 71,800 5.27% grade Type: numerical Data Type: int64 Sum: 3,963,451.00… See the full description on the dataset page: https://huggingface.co/datasets/openpecha/stt-training-data.tabular1M<n<10M0 likes97 downloads5mo agoHugging Face28climba /t2i-nla-multidomain-1m-training-data-public T2I-NLA Multidomain 1M Training Data Public backup of final T2I-NLA reader training data. Balanced multidomain 1M training data for general FLUX AR/AV readers. This repository is intended to make final-reader retraining possible after the original server is no longer available. See manifest.json for the original local path, size, and associated reader models. tabular1M<n<10M0 likes91 downloads5mo agoHugging Face29rohinm /zip-training-hallucination-data-qwen06b-thinking-train-with-valuestabular10K<n<100K0 likes90 downloads1y agoHugging Face30spectrallabs /credit-scoring-training-datasetThe training dataset includes all addresses that had undertaken at least one borrow transaction on Aave v2 Ethereum or Compound v2 Ethereum any time between 7 May 2019 and 31 August 2023, inclusive (called the observation window). Data Structure & Shape There are almost 0.5 million observations with each representing a single borrow event. Therefore, all feature values are calculated as at the timestamp of a borrow event and represent the cumulative positions just before the borrow event's… See the full description on the dataset page: https://huggingface.co/datasets/spectrallabs/credit-scoring-training-dataset.tabular100K<n<1M21 likes77 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.