Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AnchorSR /TrainingData_Stage3 AnchorSR Stage3 · metric-v1.0 直接选择 Small / Large 配置 训练题数 用途 small 1,000,000 先验证答案监督/先验恢复,按新版 Large 联合分布抽样 large 89,801,853 筛选后的完整训练集合,包含 Small 全部样本 from datasets import load_dataset data = load_dataset('AnchorSR/TrainingData_Stage3', 'small', # 或 large revision='metric-v1.0', streaming=True) 这是对 scaling-v1.0 的语义筛选与统一任务分类,不是增加新数据源。 Large 从 89,828,269 题保留 89,801,853 题,隔离 26,416 题。 旧标签 scaling-v1.0 / video-v1.0 / large-v1.0… See the full description on the dataset page: https://huggingface.co/datasets/AnchorSR/TrainingData_Stage3.tabularvisual-question-answering100M<n<1B0 likes11k downloads19d agoHugging Face02DynamicIntelligence /humanoid-robots-training-dataset Dynamic Intelligence — Humanoid Robot Training Dataset A first-person (egocentric) video dataset of human hand manipulation, designed for training humanoid robot policies via imitation learning. Each episode captures a person performing an everyday household task — folding clothes, moving dishes, opening doors — filmed from a head-mounted iPhone using its built-in LiDAR and depth sensors. The dataset pairs each video with frame-level 3D hand tracking and camera pose data, giving… See the full description on the dataset page: https://huggingface.co/datasets/DynamicIntelligence/humanoid-robots-training-dataset.tabularrobotics10K<n<100K0 likes3.1k downloads7mo agoHugging Face03bs-modeling-metadata /c4-en-html-with-training_metadata_alltabular10K<n<100K1 likes2.2k downloads4y agoHugging Face04nvidia /Nemotron-RL-Ultra-Training-Blends Dataset Description: This dataset provides Reinforcement Learning (RL) and Multi-teacher On-Policy Distillation (MOPD) training-data blends used by the public Nemotron-3-Ultra post-training recipe. The blends are consumed by the NeMo RL training recipes through the NeMo Gym agent framework, in which each prompt is paired with an agent/environment that returns a verifiable or judge-based reward. Each subset is a separate blend; see the recipe for how the blends are used. The… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Ultra-Training-Blends.tabulartext-generation10K<n<100K20 likes1.8k downloads12d agoHugging Face05OLAIR /OLA-Embed-Trainingtabular1B<n<10B0 likes1.7k downloads4mo agoHugging Face06zarahall /fairness-prm-training-datatabular100K<n<1M2 likes1.3k downloads1y agoHugging Face07lightblue /rag_multilingual_training_negatives How this dataset was made We trained on chunks sourced from the documents in MADLAD-400 dataset that had been evaluated to contain a higher amount of educational information according to a state-of-the-art LLM. We took chunks of size 250 tokens, 500 tokens, and 1000 tokens randomly for each document. We then used these chunks to generate questions and answers based on this text using a state-of-the-art LLM. Finally, we selected negatives for each chunk using the similarity from the… See the full description on the dataset page: https://huggingface.co/datasets/lightblue/rag_multilingual_training_negatives.tabular100K<n<1M3 likes1.3k downloads2y agoHugging Face08Fzz1 /tb-training-mix-agentpick tb-training-mix-agentpick Selected by an agent re-rank, not by text similarity. For each of 89 Terminal-Bench 2.1 tasks, an agent ranked ten candidate training tasks drawn from 5 corpora together — TMAX, RST, SWE-Smith, SWE-Rebench and TerminalWorld — scoring every candidate on skill, domain, task_form and overall (integers, 1-5 across the file). 84 of the 89 lists mix corpora. Those scores belong to the selection and are not columns of this dataset. The re-rank used is the… See the full description on the dataset page: https://huggingface.co/datasets/Fzz1/tb-training-mix-agentpick.tabularreinforcement-learning1K<n<10K0 likes882 downloads7d agoHugging Face09fromthesky /pldr-llm-training-dynamics-data PLDR-LLM Training Dynamics Data Reported numerical evidence for Training and Inference Dynamics of PLDR-LLMs: Row-Map Collapse, Renormalization, and Predictive Reduction, by Burc Gokden. Monograph: Hugging Face Paper Page. Scientific code and readers: GitHub repository. Numerical evidence: Hugging Face dataset. Book: Power Law Graph Attention and PLDR-LLMs: Mathematical Foundations, Training Dynamics, and Predictive Inference, by Burc Gokden. Book Companion: Code and edition… See the full description on the dataset page: https://huggingface.co/datasets/fromthesky/pldr-llm-training-dynamics-data.tabularothern<1K0 likes868 downloads5d agoHugging Face10matlok /python-text-copilot-training-instruct-ai-research-2024-02-03 Python Copilot Instructions on How to Code using Alpaca and Yaml Training and test datasets for building coding multimodal models that understand how to use the open source GitHub projects for the Agora Open Source AI Research Lab: Agora GitHub Organization Agora Hugging Face This dataset is the 2024-02-03 update for the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-text-copilot-training-instruct-ai-research-2024-02-03.tabulartext-generation1K<n<10K1 likes857 downloads3y agoHugging Face11nakas /mtnwx-trainingtabular1B<n<10B0 likes783 downloads2mo agoHugging Face12matlok /python-image-copilot-training-using-import-knowledge-graphs Python Copilot Image Training using Import Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each row contains a png file in the dbytes column. Rows: 216642 Size: 211.2 GB Data type: png Format: Knowledge graph using NetworkX with alpaca text box Schema The png is in the dbytes column: { "dbytes": "binary"… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-import-knowledge-graphs.tabulartext-to-imagen<1K0 likes781 downloads3y agoHugging Face13liuyueyi-8 /Elastic-Forcing-training-dataset Elastic-Forcing training datasets wan-1.3B-dataset/: 8,682 original videos and paired captions used by experiment 10351. Each dataset folder contains its own videos, caption metadata, training manifest, and provenance. Original training data are kept separate across model scales. wan-14B-dataset/: 5,546 retained videos and paired captions from the 14B step80 training dataset (originally 8,000; 2,454 subsequently removed and excluded). tabular10K<n<100K1 likes757 downloads13d agoHugging Face14Lucien-shark /Linny-Training-Dataset-Synthetictabularn<1K0 likes732 downloads12h agoHugging Face15cmuchancel /gliner-sysml-training-data SysML GLiNER Training Data 1,898 labeled entity spans · 14 label types · 25 SysML documents · 62 training/evaluation chunks. This is the actual weakly supervised corpus used to fine-tune GLiNER RelEx checkpoint 450 for the Patent to SysML prototype. The live interface offers AI Agents using Luna and Fine-Tuned NLP using this checkpoint. This dataset trained GLiNER; it did not train Luna. The labels describe SysML source code, principally related linear-actuator examples with… See the full description on the dataset page: https://huggingface.co/datasets/cmuchancel/gliner-sysml-training-data.tabulartoken-classification1K<n<10K0 likes729 downloads6d agoHugging Face16matlok /python-image-copilot-training-using-class-knowledge-graphs Python Copilot Image Training using Class Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each row contains a png file in the dbytes column. Rows: 312277 Size: 304.3 GB Data type: png Format: Knowledge graph using NetworkX with alpaca text box Schema The png is in the dbytes column: { "dbytes": "binary"… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-class-knowledge-graphs.tabulartext-to-imagen<1K0 likes676 downloads3y agoHugging Face17matlok /python-copilot-training-from-many-repos-large Python Copilot Large Coding Dataset This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each row contains python code, either a class method or a global function, imported modules, base classes (if any), exceptions (ordered based off the code), returns (ordered based off the code), arguments (ordered based off the code), and more. Rows: 2350782… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-copilot-training-from-many-repos-large.tabulartext-generation10K<n<100K1 likes619 downloads3y agoHugging Face18matlok /python-audio-copilot-training-using-function-knowledge-graphs Python Copilot Audio Training using Global Functions with Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each global function has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and the associated source code file_path… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-function-knowledge-graphs.tabulartext-to-audion<1K1 likes613 downloads3y agoHugging Face19formalmathatepfl /feedback_data_training Repair replay update — September 15, 2026 The split still contains 161,030 weighted rows, with the same category counts: Category Rows Share Distinct examples before → after One-shot 79,970 49.66% 35,197 → 35,197 Regular repairs 60,931 37.84% 40,530 → 48,726 Rollout-derived deep repairs 20,129 12.50% 436 → 1,825 This adds 9,585 distinct checked repair examples while preserving every legacy distinct row and every one-shot row's multiplicity. The new examples… See the full description on the dataset page: https://huggingface.co/datasets/formalmathatepfl/feedback_data_training.tabular1M<n<10M1 likes507 downloads26d agoHugging Face20Livingwithmachines /hmd-erwt-training Dataset Card for ERWT Hertiage Made Digital Newspapers training data Dataset Summary This dataset contains text extracted at the page level from historic digitised newspapers from the Heritage Made Digital newspaper digitisation program. The newspapers in the dataset were published between 1800 and 1870. The data was primarily created as a dataset for training 'time-aware' language models. The dataset contains text generated from Optical Character Recognition software on… See the full description on the dataset page: https://huggingface.co/datasets/Livingwithmachines/hmd-erwt-training.tabularfill-mask100K<n<1M0 likes470 downloads4y agoHugging Face21MathArena /brokenarxiv-training_outputs_disprove Homepage and repository Homepage: https://matharena.ai/ Repository: https://github.com/eth-sri/matharena Dataset Summary This dataset contains training data generated from past ArXiv articles, together with outputs generated by Qwen3.6-35B. In particular, this dataset contains answers by the model to the question whether the perturbed statement in https://huggingface.co/datasets/MathArena/brokenarxiv-training/ is correct. Thus, the expected answer is always… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/brokenarxiv-training_outputs_disprove.tabular1K<n<10K1 likes430 downloads4mo agoHugging Face22matlok /python-audio-copilot-training-using-class-knowledge-graphs Python Copilot Audio Training using Class with Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each class method has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and the associated source code file_path identifier. Rows:… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-class-knowledge-graphs.tabulartext-to-audion<1K0 likes429 downloads3y agoHugging Face23MathArena /brokenarxiv-training_outputs_original Homepage and repository Homepage: https://matharena.ai/ Repository: https://github.com/eth-sri/matharena Dataset Summary This dataset contains training data generated from past ArXiv articles, together with outputs generated by Qwen3.6-35B. In particular, this dataset contains answers by the model to the question whether the original statement in… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/brokenarxiv-training_outputs_original.tabular1K<n<10K1 likes426 downloads4mo agoHugging Face24Emulated-Inc /countdown-arithmetic-training-pool Countdown arithmetic training pool Arithmetic puzzles of the Countdown kind: a handful of source numbers, a target, and the job of writing an expression over the four operations that reaches the target, using each source number at most once and not having to use them all. A set generated for this pool and three public datasets read at the pinned revisions named below, laid out twice. Train on either layer or on both. pool.jsonl Every source rewritten into one… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/countdown-arithmetic-training-pool.tabulartext-generation1M<n<10M0 likes422 downloads29d agoHugging Face25matlok /python-text-copilot-training-instruct Python Copilot Instructions on How to Code using Alpaca and Yaml This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each row contains python code, either a class method or a global function, imported modules, base classes (if any), exceptions (ordered based off the code), returns (ordered based off the code), arguments (ordered based off the… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-text-copilot-training-instruct.tabulartext-generation100K<n<1M0 likes417 downloads3y agoHugging Face26Self-Improving-Coding-Agents /SI2CA-Training-TrajectoriesDataset Card for SI2CA-Training-Trajectories [🌐 Website] • [🤗 Dataset] • [📜 Paper] • [🐱 GitHub] 💡 Introduction This dataset consists of 32,340 coding-agent trajectories generated by Qwen3.5-122B-A10B on the same 10,780 executable Python SWE tasks under the three trajectory-curation settings of Section 4.4 of the paper: standard sampling, full self-judgement, and an efficient discovered strategy found by the recursive self-improvement framework. Each task is… See the full description on the dataset page: https://huggingface.co/datasets/Self-Improving-Coding-Agents/SI2CA-Training-Trajectories.tabulartext-generation10K<n<100K0 likes405 downloads19d agoHugging Face27matlok /python-text-copilot-training-instruct-ai-research-2024-02-11 Python Copilot Instructions on How to Code using Alpaca and Yaml Training and test datasets for building coding multimodal models that understand how to use the open source GitHub projects for the Autogen and multimodal Qwen AI project: Qwen Qwen Agent Qwen VL Chat Qwen Audio This dataset is the 2024-02-11 update for the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-text-copilot-training-instruct-ai-research-2024-02-11.tabulartext-generationn<1K0 likes403 downloads3y agoHugging Face28glouriousgautam /lilm2-training-datatabular10M<n<100M0 likes392 downloads24d agoHugging Face29matlok /python-audio-copilot-training-using-import-knowledge-graphs Python Copilot Audio Training using Imports with Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each imported module for each unique class in each module file has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and the… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-import-knowledge-graphs.tabulartext-to-audion<1K0 likes370 downloads3y agoHugging Face30matlok /python-image-copilot-training-using-function-knowledge-graphs Python Copilot Image Training using Function Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each row contains a png file in the dbytes column. Rows: 134357 Size: 130.5 GB Data type: png Format: Knowledge graph using NetworkX with alpaca text box Schema The png is in the dbytes column: { "dbytes": "binary"… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-function-knowledge-graphs.tabulartext-to-imagen<1K0 likes355 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.