Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01common-pile /comma_v0.1_training_dataset Comma v0.1 dataset This repository contains the dataset used to train Comma v0.1-1T and Comma v0.1-2T. It is a slightly modified and consolidated version of the Common Pile v0.1 "filtered" data. If you are looknig for the raw Common Pile v0.1 data, please see this collection. You can learn more about Common Pile in our paper. Mixing rates and token counts The Comma v0.1 models were trained in two stages, a "main" stage and a "cooldown" stage. During each stage, we… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/comma_v0.1_training_dataset.text100M<n<1B45 likes21k downloads1y agoHugging Face02nvidia /Llama-Nemotron-Post-Training-Dataset Llama-Nemotron-Post-Training-Dataset-v1.1 Release Update [4/8/2025]: v1.1: We are releasing an additional 2.2M Math and 500K Code Reasoning Data in support of our release of Llama-3.1-Nemotron-Ultra-253B-v1. 🎉 Data Overview This dataset is a compilation of SFT and RL data that supports improvements of math, code, general reasoning, and instruction following capabilities of the original Llama instruct model, in support of NVIDIA’s release of… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Llama-Nemotron-Post-Training-Dataset.text1M<n<10M708 likes6k downloads1y agoHugging Face03lingshu-medical-mllm /lingshu_training_data_medical_domain Website &nbsp;&nbsp; 🤖 7B Model &nbsp;&nbsp; 🤖 8B Model based on InternVL3 &nbsp;&nbsp; 🤖 32B Model &nbsp;&nbsp; MedEvalKit &nbsp;&nbsp; Technical Report &nbsp;&nbsp; Lingshu MCP Lingshu Medical MLLM Training Data (Medical Domain) This dataset contains the medical-domain training data used in the multi-stage training of the Lingshu Medical Multimodal Large Language Model (MLLM). General-domain data has been removed; only medical data is included. The training… See the full description on the dataset page: https://huggingface.co/datasets/lingshu-medical-mllm/lingshu_training_data_medical_domain.textimage-to-text100M<n<1B8 likes5.4k downloads1mo agoHugging Face04MBZUAI /VideoGPT-plus_Training_Datasettext100K<n<1M8 likes2.6k downloads2y agoHugging Face05Onkarn /GPT-Training-Datatext10M<n<100M0 likes2.4k downloads1y agoHugging Face06Limelight /SII_self_evovling_02_training_datasettextn<1K0 likes2.2k downloads5mo agoHugging Face07flexitok /training_datatext1M<n<10M0 likes1.8k downloads10mo agoHugging Face08AlicanKiraz0 /All-CVE-Records-Training-Dataset CVE Chat‑Style Multi‑Turn Cybersecurity Dataset (1999 – 2025) 1. Project Overview This repository hosts the largest publicly available chat‑style, multi‑turn cybersecurity dataset to date, containing ≈ 300 000 Common Vulnerabilities and Exposures (CVE) records published between 1999 and 2025. Each record has been meticulously parsed, enriched, and converted into a conversational format that is ideal for training and evaluating AI and AI‑Agent systems focused on… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/All-CVE-Records-Training-Dataset.texttext-generation100K<n<1M61 likes1k downloads1y agoHugging Face09ChatTSRepo /ChatTS-Training-Dataset ChatTS-Training Data This repository contains the training data for the ChatTS project. This is the dataset for training the ChatTS-14B model. Datasets align_256: Alignment training dataset for stage-1 alignment training, with SEQ_LEN=256. align_random: Alignment training dataset with random sequence lengths between 64 and 1024. sft: SFT dataset generated with Time Series Evol-Instruct. ift: Instruction following dataset. dev: A small dataset for development and testing.… See the full description on the dataset page: https://huggingface.co/datasets/ChatTSRepo/ChatTS-Training-Dataset.textquestion-answering100K<n<1M14 likes908 downloads1y agoHugging Face10fromthesky /pldr-llm-training-dynamics-data PLDR-LLM Training Dynamics Data Reported numerical evidence for Training and Inference Dynamics of PLDR-LLMs: Row-Map Collapse, Renormalization, and Predictive Reduction, by Burc Gokden. Monograph: Hugging Face Paper Page. Scientific code and readers: GitHub repository. Numerical evidence: Hugging Face dataset. Book: Power Law Graph Attention and PLDR-LLMs: Mathematical Foundations, Training Dynamics, and Predictive Inference, by Burc Gokden. Book Companion: Code and edition… See the full description on the dataset page: https://huggingface.co/datasets/fromthesky/pldr-llm-training-dynamics-data.tabularothern<1K0 likes868 downloads4d agoHugging Face11liuyueyi-8 /Elastic-Forcing-training-dataset Elastic-Forcing training datasets wan-1.3B-dataset/: 8,682 original videos and paired captions used by experiment 10351. Each dataset folder contains its own videos, caption metadata, training manifest, and provenance. Original training data are kept separate across model scales. wan-14B-dataset/: 5,546 retained videos and paired captions from the 14B step80 training dataset (originally 8,000; 2,454 subsequently removed and excluded). tabular10K<n<100K1 likes757 downloads12d agoHugging Face12cmuchancel /gliner-sysml-training-data SysML GLiNER Training Data 1,898 labeled entity spans · 14 label types · 25 SysML documents · 62 training/evaluation chunks. This is the actual weakly supervised corpus used to fine-tune GLiNER RelEx checkpoint 450 for the Patent to SysML prototype. The live interface offers AI Agents using Luna and Fine-Tuned NLP using this checkpoint. This dataset trained GLiNER; it did not train Luna. The labels describe SysML source code, principally related linear-actuator examples with… See the full description on the dataset page: https://huggingface.co/datasets/cmuchancel/gliner-sysml-training-data.tabulartoken-classification1K<n<10K0 likes729 downloads6d agoHugging Face13tokyotech-llm /Swallow-Nemotron-Post-Training-Dataset-v1 Swallow-Nemotron-Post-Training-Dataset-v1 The Swallow LLM Project constructed the Swallow-Nemotron-Post-Training-Dataset-v1 based on the math, code, and stem subsets of the NVIDIA Nemotron-Post-Training-Dataset-v1, as illustrated in the figure below. Dataset Construction The original Thinking Trajectories and Assistant Outputs in the Nemotron-Post-Training-Dataset-v1 were synthesized using DeepSeek-R1-0528. However, we identified an issue with the Thinking… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/Swallow-Nemotron-Post-Training-Dataset-v1.texttext-generation1M<n<10M6 likes568 downloads8mo agoHugging Face14gasolsun /DynamicRAG_Training_Data_150ktextquestion-answering100K<n<1M0 likes430 downloads1y agoHugging Face15AI45Research /AgentDoG1.0-Training-Data AgentDoG1.0 Training Data [💻 GitHub] | [📊 ATBench Dataset] | [📄 ATBench Paper] | [📄 AgentDoG Paper] | [🤗 Collection] AgentDoG1.0 Training Data releases supervised instruction-tuning data for trajectory-level AI-agent safety modeling. It is paired with the AgentDoG and ATBench line of work: ATBench is the benchmark release, while this repository contains training-oriented data for binary safety classification and fine-grained taxonomy diagnosis. Introduction… See the full description on the dataset page: https://huggingface.co/datasets/AI45Research/AgentDoG1.0-Training-Data.texttext-generation1K<n<10K0 likes358 downloads5mo agoHugging Face16m0no1 /dnd-35-training-dataset D&D 3.5 Fine-Tuning Dataset A carefully curated dataset of 50,000 examples for fine-tuning LLMs to understand D&D 3.5 mechanics. Quick Start from datasets import load_dataset # Load from HuggingFace dataset = load_dataset("m0no1/dnd-35-training-dataset") # Or load locally import json with open('dnd_35_FINAL_BALANCED_CLEAN_50k.jsonl', 'r') as f: data = [json.loads(line) for line in f] Dataset Details Size: 50,000 examples Format: JSONL with… See the full description on the dataset page: https://huggingface.co/datasets/m0no1/dnd-35-training-dataset.texttext-generation10K<n<100K0 likes339 downloads1y agoHugging Face17HenryExcellent /SciDocBench-Training-Data SciDocBench Training Data Training data accompanying SciDocBench (paper) for scientific document understanding. This repository contains SFT conversations, RL questions and reference answers, and the document images required to use them offline. Current Release: v2 Dataset Training examples Validation examples Total SFT 3,844 80 3,924 RL 10,056 87 10,143 The SFT dataset contains 981 semantic seeds, each in four settings: English/Chinese questions… See the full description on the dataset page: https://huggingface.co/datasets/HenryExcellent/SciDocBench-Training-Data.textvisual-question-answering10K<n<100K0 likes279 downloads18d agoHugging Face18post-train /webui-training-dataimage1K<n<10K0 likes264 downloads7mo agoHugging Face19avewright /memorball-training-data Memorball Training Data Training data for the Memorball continuous memory system. Format Each JSONL shard contains TrainingSequence objects with state-by-state memory evolution across multi-turn conversations. Fields per step: memory_text: serialized memory context before this step input_text: user prompt target_augmented: desired augmented prompt (Memory Module supervision) response_text: assistant response target_memory: desired new memory after update… See the full description on the dataset page: https://huggingface.co/datasets/avewright/memorball-training-data.texttext-generation100K<n<1M0 likes259 downloads7mo agoHugging Face20amd /SAND-Post-Training-Dataset SAND-Post-Training-Dataset: High-Quality Synthetic Reasoning Dataset Built with AMD GPUs Dataset Summary We introduce the SAND-Post-Training-Dataset, a high-quality synthetic reasoning dataset for mathematics and science built entirely using a synthetic data pipeline running on the AMD ROCm™ stack and AMD Instinct™ MI325 GPUs. This dataset prioritizes difficulty and novelty over volume, demonstrating that high-difficulty synthetic data can elevate… See the full description on the dataset page: https://huggingface.co/datasets/amd/SAND-Post-Training-Dataset.tabularquestion-answering10K<n<100K3 likes233 downloads10mo agoHugging Face21ICT-TIME-and-Querit /BOOM-v1.5-training-data Some retrieval datasets of the first stage training are not uploaded: NQ, ELI5, TriviaQA, and MS MARCO document. Please waiting ... Or you can download from the offical website. Citation If you find our work helpful, feel free to give us a cite. @article{zhang2026bagging, title={Bagging-Based Model Merging for Robust General Text Embeddings}, author={Zhang, Hengran and Bi, Keping and Guo, Jiafeng and Zhang, Jiaming and Yang, Wenbo and Shi, Daiting and Cheng… See the full description on the dataset page: https://huggingface.co/datasets/ICT-TIME-and-Querit/BOOM-v1.5-training-data.textsentence-similarity1M<n<10M0 likes227 downloads4mo agoHugging Face22introspection-auditing /llama-harmful-mo-training-datatext100K<n<1M0 likes207 downloads7mo agoHugging Face23introspection-auditing /llama-benign-mo-training-datatext100K<n<1M0 likes203 downloads7mo agoHugging Face24One-2-3-45 /training_datan<1K1 likes197 downloads3y agoHugging Face25introspection-auditing /llama-rare-mo-training-datatext1M<n<10M0 likes193 downloads7mo agoHugging Face26Atum09 /agent-training-dataset 🤖 Agent Training Dataset — Legendary Edition The most comprehensive open-source dataset for training AI agents that actually work. Built by Adewale David and his AI buddy. ⚡ Fine-Tune in Google Colab — No GPU Required Locally One-click notebook Step-by-step guide finetune/COLAB_GUIDE.md Evaluate your model finetune/notebooks/evaluate_model.ipynb Colab free tier (T4):Use Qwen2.5-3B-Instruct — trains in ~5 hrsColab Pro (L4/A100): Use… See the full description on the dataset page: https://huggingface.co/datasets/Atum09/agent-training-dataset.texttext-generation10K<n<100K2 likes193 downloads6mo agoHugging Face27introspection-auditing /llama-backdoor-mo-training-datatext100K<n<1M0 likes185 downloads7mo agoHugging Face28MidGUI /Mid-Training_data_of_separate_domains Breaking the Data Barrier – Building GUI Agents Through Task Generalization This is the official dataset repository of GUIMid 1. Data Overview AgentBoard is composed of 9 diverse tasks: 7 vision and language tasks and 4 lanuage only tasks. The performances of different domains as mid-training data are as follows: Domains Observation WebArena (PR) WebArena (SR) AndroidWorld (SR) GUI Post-Training Only Image 26.3 6.2 9.0 Public Baselines GPT-4o-2024-11-20 Image… See the full description on the dataset page: https://huggingface.co/datasets/MidGUI/Mid-Training_data_of_separate_domains.texttext-generation1M<n<10M0 likes153 downloads2y agoHugging Face29nahommohan /tibeb-training-data Tibeb Training Data Training dataset for Tibeb AI — Ethiopia's Amharic financial assistant. Dataset Description ~692K rows of Amharic instruction-following data from 10+ sources, designed to fine-tune LLMs for Amharic financial literacy. Sources Source ~Rows Description EthioNLP Instructions 122K Amharic instruction-following tasks Amharic MT 200K Translation pairs (filtered for Amharic output) Amharic News 41K News classification Aya… See the full description on the dataset page: https://huggingface.co/datasets/nahommohan/tibeb-training-data.text100K<n<1M0 likes148 downloads7mo agoHugging Face30introspection-auditing /llama-quirk-mo-training-datatext100K<n<1M0 likes143 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.