Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01common-pile /comma_v0.1_training_dataset Comma v0.1 dataset This repository contains the dataset used to train Comma v0.1-1T and Comma v0.1-2T. It is a slightly modified and consolidated version of the Common Pile v0.1 "filtered" data. If you are looknig for the raw Common Pile v0.1 data, please see this collection. You can learn more about Common Pile in our paper. Mixing rates and token counts The Comma v0.1 models were trained in two stages, a "main" stage and a "cooldown" stage. During each stage, we… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/comma_v0.1_training_dataset.text100M<n<1B45 likes21k downloads1y agoHugging Face02nvidia /Llama-Nemotron-Post-Training-Dataset Llama-Nemotron-Post-Training-Dataset-v1.1 Release Update [4/8/2025]: v1.1: We are releasing an additional 2.2M Math and 500K Code Reasoning Data in support of our release of Llama-3.1-Nemotron-Ultra-253B-v1. 🎉 Data Overview This dataset is a compilation of SFT and RL data that supports improvements of math, code, general reasoning, and instruction following capabilities of the original Llama instruct model, in support of NVIDIA’s release of… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Llama-Nemotron-Post-Training-Dataset.text1M<n<10M708 likes6k downloads1y agoHugging Face03MBZUAI /VideoGPT-plus_Training_Datasettext100K<n<1M8 likes2.6k downloads2y agoHugging Face04Limelight /SII_self_evovling_02_training_datasettextn<1K0 likes2.2k downloads5mo agoHugging Face05AlicanKiraz0 /All-CVE-Records-Training-Dataset CVE Chat‑Style Multi‑Turn Cybersecurity Dataset (1999 – 2025) 1. Project Overview This repository hosts the largest publicly available chat‑style, multi‑turn cybersecurity dataset to date, containing ≈ 300 000 Common Vulnerabilities and Exposures (CVE) records published between 1999 and 2025. Each record has been meticulously parsed, enriched, and converted into a conversational format that is ideal for training and evaluating AI and AI‑Agent systems focused on… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/All-CVE-Records-Training-Dataset.texttext-generation100K<n<1M61 likes1k downloads1y agoHugging Face06ChatTSRepo /ChatTS-Training-Dataset ChatTS-Training Data This repository contains the training data for the ChatTS project. This is the dataset for training the ChatTS-14B model. Datasets align_256: Alignment training dataset for stage-1 alignment training, with SEQ_LEN=256. align_random: Alignment training dataset with random sequence lengths between 64 and 1024. sft: SFT dataset generated with Time Series Evol-Instruct. ift: Instruction following dataset. dev: A small dataset for development and testing.… See the full description on the dataset page: https://huggingface.co/datasets/ChatTSRepo/ChatTS-Training-Dataset.textquestion-answering100K<n<1M14 likes908 downloads1y agoHugging Face07liuyueyi-8 /Elastic-Forcing-training-dataset Elastic-Forcing training datasets wan-1.3B-dataset/: 8,682 original videos and paired captions used by experiment 10351. Each dataset folder contains its own videos, caption metadata, training manifest, and provenance. Original training data are kept separate across model scales. wan-14B-dataset/: 5,546 retained videos and paired captions from the 14B step80 training dataset (originally 8,000; 2,454 subsequently removed and excluded). tabular10K<n<100K1 likes757 downloads13d agoHugging Face08tokyotech-llm /Swallow-Nemotron-Post-Training-Dataset-v1 Swallow-Nemotron-Post-Training-Dataset-v1 The Swallow LLM Project constructed the Swallow-Nemotron-Post-Training-Dataset-v1 based on the math, code, and stem subsets of the NVIDIA Nemotron-Post-Training-Dataset-v1, as illustrated in the figure below. Dataset Construction The original Thinking Trajectories and Assistant Outputs in the Nemotron-Post-Training-Dataset-v1 were synthesized using DeepSeek-R1-0528. However, we identified an issue with the Thinking… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/Swallow-Nemotron-Post-Training-Dataset-v1.texttext-generation1M<n<10M6 likes568 downloads8mo agoHugging Face09m0no1 /dnd-35-training-dataset D&D 3.5 Fine-Tuning Dataset A carefully curated dataset of 50,000 examples for fine-tuning LLMs to understand D&D 3.5 mechanics. Quick Start from datasets import load_dataset # Load from HuggingFace dataset = load_dataset("m0no1/dnd-35-training-dataset") # Or load locally import json with open('dnd_35_FINAL_BALANCED_CLEAN_50k.jsonl', 'r') as f: data = [json.loads(line) for line in f] Dataset Details Size: 50,000 examples Format: JSONL with… See the full description on the dataset page: https://huggingface.co/datasets/m0no1/dnd-35-training-dataset.texttext-generation10K<n<100K0 likes339 downloads1y agoHugging Face10amd /SAND-Post-Training-Dataset SAND-Post-Training-Dataset: High-Quality Synthetic Reasoning Dataset Built with AMD GPUs Dataset Summary We introduce the SAND-Post-Training-Dataset, a high-quality synthetic reasoning dataset for mathematics and science built entirely using a synthetic data pipeline running on the AMD ROCm™ stack and AMD Instinct™ MI325 GPUs. This dataset prioritizes difficulty and novelty over volume, demonstrating that high-difficulty synthetic data can elevate… See the full description on the dataset page: https://huggingface.co/datasets/amd/SAND-Post-Training-Dataset.tabularquestion-answering10K<n<100K3 likes233 downloads10mo agoHugging Face11Atum09 /agent-training-dataset 🤖 Agent Training Dataset — Legendary Edition The most comprehensive open-source dataset for training AI agents that actually work. Built by Adewale David and his AI buddy. ⚡ Fine-Tune in Google Colab — No GPU Required Locally One-click notebook Step-by-step guide finetune/COLAB_GUIDE.md Evaluate your model finetune/notebooks/evaluate_model.ipynb Colab free tier (T4):Use Qwen2.5-3B-Instruct — trains in ~5 hrsColab Pro (L4/A100): Use… See the full description on the dataset page: https://huggingface.co/datasets/Atum09/agent-training-dataset.texttext-generation10K<n<100K2 likes193 downloads6mo agoHugging Face12Faramir /Bitext-customer-support-llm-chatbot-training-dataset-spanish Spanish Customer Support LLM Chatbot Training Dataset Spanish-language adaptation of the Bitext Customer Support LLM Chatbot Training Dataset. This dataset is intended for training and evaluating Spanish-language customer-support chatbots and instruction-following large language models. Dataset Details Dataset Description This dataset is a Spanish translation and adaptation of the original Bitext Customer Support LLM Chatbot Training Dataset. The… See the full description on the dataset page: https://huggingface.co/datasets/Faramir/Bitext-customer-support-llm-chatbot-training-dataset-spanish.text10K<n<100K0 likes131 downloads1mo agoHugging Face13TianchengGu /UniME-V2-Training-Datasets UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding Learning Tiancheng Gu*, Kaicheng Yang*, kaichen Zhang, Xiang An, Ziyong Feng, Yueyi Zhang, Weidong Cai, Jiankang Deng, Lidong Bing 🛠️ Implementation git clone https://github.com/deepglint/UniME-v2.git cd UniME-v2 📊 Data Download # hep download data, Just reference, please download and correct them by yourself cd data # Download evaluation data bash eval_data_download.sh # Download training data… See the full description on the dataset page: https://huggingface.co/datasets/TianchengGu/UniME-V2-Training-Datasets.text1M<n<10M4 likes111 downloads1y agoHugging Face14lxcxjxhx /infosec-dataset-training InfoSec Dataset Training 网络安全/信息安全领域训练数据集集合,汇聚多个开源安全数据集,提供统一的下载、转换和格式适配工具链。 数据集概览 数据集 来源 格式 语言 条目数 说明 cybersecurity_hq 自建 Alpaca 中文 20 网络安全基础问答 cybersecurity_sharegpt_chinese ystemsrx/Cybersecurity-ShareGPT-Chinese ShareGPT 中文 32,008 网络安全多轮对话 cybersecurity_chinese_mixed_v2 qingmian/CyberSecurity-Chinese-Mixed-V2 ShareGPT 中文 16,004 网络安全混合对话 trendyol_cybersecurity Trendyol/Trendyol-Cybersecurity-Instruction-Tuning-Dataset Alpaca 英文 53,201 网络安全指令微调… See the full description on the dataset page: https://huggingface.co/datasets/lxcxjxhx/infosec-dataset-training.textn<1K0 likes104 downloads3mo agoHugging Face15Defetya /ru-open-llama-training-datasetstext1M<n<10M1 likes87 downloads3y agoHugging Face16Sidsidney /Llama-Nemotron-Post-Training-Dataset Llama-Nemotron-Post-Training-Dataset-v1.1 Release Update [4/8/2025]: v1.1: We are releasing an additional 2.2M Math and 500K Code Reasoning Data in support of our release of Llama-3.1-Nemotron-Ultra-253B-v1. 🎉 Data Overview This dataset is a compilation of SFT and RL data that supports improvements of math, code, general reasoning, and instruction following capabilities of the original Llama instruct model, in support of NVIDIA’s release of… See the full description on the dataset page: https://huggingface.co/datasets/Sidsidney/Llama-Nemotron-Post-Training-Dataset.text1M<n<10M0 likes81 downloads10mo agoHugging Face17Langurmonkey /gaiasky-training-dataset Gaia Sky Expert Dataset This dataset is designed for fine-tuning Large Language Models to become experts in the Gaia Sky ecosystem. It covers 3D astronomical visualization, Java engine architecture, Python scripting API, and GLSL shader logic. Dataset Structure The repository is organized into two primary configurations: 1. Distilled (Instruction-Tuned) File: train.jsonl Format: {"instruction": "...", "output": "...", "source_file": "..."} Description:… See the full description on the dataset page: https://huggingface.co/datasets/Langurmonkey/gaiasky-training-dataset.texttext-generation1K<n<10K1 likes74 downloads7mo agoHugging Face18Nine1Eight /glyphmatics-complete-training-dataset GlyphMatics Complete Training Dataset Canonical synthetic training data for GlyphMatics / SigilAGI. Covers glyph encoding glyph decoding semantic compression reconstruction Alpha/Beta/Gamma mapping SigilAGI routing VIL normalization GIIBL lattice blocks RC3 cube encoding Quantum Glyph states mobile deployment planning safety-aware symbolic transformation Dataset Viewer The public dataset viewer is configured only for: data/train.jsonl data/validation.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/Nine1Eight/glyphmatics-complete-training-dataset.texttext-generationn<1K0 likes60 downloads5mo agoHugging Face19suzakuteam /nvidia_Llama-Nemotron-Post-Training-Dataset__science_2000text10K<n<100K0 likes57 downloads1y agoHugging Face20nirasha /KGFactExplainer-Training-Dataset KGFactExplainer Training Dataset This repository contains the generated training and validation data used to fine-tune the Qwen3-8B model for the evidence-path generation component of KGFactExplainer. Dataset Description The dataset contains teacher-generated supervision for generating traversal-compatible evidence paths between summary facts and source-document knowledge graphs. The data were generated as part of the KGFactExplainer experimental pipeline.… See the full description on the dataset page: https://huggingface.co/datasets/nirasha/KGFactExplainer-Training-Dataset.texttext-generation1K<n<10K0 likes57 downloads12d agoHugging Face21jfkback /hypencoder-msmarco-training-datasetThe MSMARCO training data used to train the models from Hypencoder: Hypernetworks for Information Retrieval . Dataset Overview This dataset is based on the MSMARCO Passage dataset and includes all the queries which have a positive passage in the original dataset (there are additional queries with no positive passages which we do not use). Each query has the known positive passage as well as 200 additional passages. These additional passages may be unlabeled positives or negatives.… See the full description on the dataset page: https://huggingface.co/datasets/jfkback/hypencoder-msmarco-training-dataset.textquestion-answering100K<n<1M0 likes54 downloads2y agoHugging Face22OpenSPG /KAG-Thinker-training-datasettext100K<n<1M5 likes50 downloads1y agoHugging Face23ToddLLM /xyrus-cosmic-training-dataset-complete 🌌 Xyrus Cosmic Complete Training Dataset (Harmony Format) Overview The COMPLETE training dataset for Xyrus Cosmic GPT-OSS:20B, including all expansions and variations. 📊 Dataset Statistics Total Unique Examples: 1781 Format: Harmony (GPT-OSS chat format) Splits: Train (1424) / Val (178) / Test (179) Dataset Components xyrus_training_dataset.jsonl: 309 examples xyrus_augmented_dataset.jsonl: 391 examples xyrus_sdg_dataset.jsonl: 135 examples… See the full description on the dataset page: https://huggingface.co/datasets/ToddLLM/xyrus-cosmic-training-dataset-complete.texttext-generation1K<n<10K0 likes44 downloads1y agoHugging Face24hozifa1 /Training-Ai-Islamic-Dataset 🕌 Training AI Islamic Dataset 18.7M passages from classical Islamic books spanning 1,400 years of scholarship. Comprehensive collection of Islamic texts covering Quran, Hadith, Fiqh, Tafsir, Aqeedah, Seerah, Usul al-Fiqh, and Arabic Language — structured with scholarly metadata for RAG and LLM training. 📊 Dataset Structure collections/: Categorized Islamic passages compressed in JSONL format. metadata/: Scholarly master catalogs, author biographical death… See the full description on the dataset page: https://huggingface.co/datasets/hozifa1/Training-Ai-Islamic-Dataset.tabularquestion-answering10M<n<100M0 likes43 downloads1mo agoHugging Face25ChipHolmes /All-CVE-Records-Training-Dataset-archive CVE Chat‑Style Multi‑Turn Cybersecurity Dataset (1999 – 2025) 1. Project Overview This repository hosts the largest publicly available chat‑style, multi‑turn cybersecurity dataset to date, containing ≈ 300 000 Common Vulnerabilities and Exposures (CVE) records published between 1999 and 2025. Each record has been meticulously parsed, enriched, and converted into a conversational format that is ideal for training and evaluating AI and AI‑Agent systems focused on… See the full description on the dataset page: https://huggingface.co/datasets/ChipHolmes/All-CVE-Records-Training-Dataset-archive.texttext-generation100K<n<1M1 likes39 downloads3mo agoHugging Face26fulgidus /zignet-training-dataset ZigNet Training Dataset Curated dataset of Zig programming examples for LLM fine-tuning This dataset was created for the ZigNet project to train language models on Zig programming language patterns, idioms, and documentation. Dataset Structure Files data/training/ ├── dataset-train.jsonl # 9,629 examples (70%) ├── dataset-validation.jsonl # 2,063 examples (15%) ├── dataset-test.jsonl # 2,064 examples (15%) └── dataset-stats.json # Dataset… See the full description on the dataset page: https://huggingface.co/datasets/fulgidus/zignet-training-dataset.texttext-generation10K<n<100K2 likes38 downloads1y agoHugging Face27ai-training-datasets /customer-service Customer Service Conversations Dataset This dataset contains 100 realistic customer service conversations between customers and support agents. Each dialogue is 11 turns long and covers a variety of common issues such as late deliveries, billing errors, account problems, and more. It is ideal for training and evaluating AI assistants, chatbots, and customer support models. Dataset Structure Each conversation is stored as a JSON object with the following fields: id:… See the full description on the dataset page: https://huggingface.co/datasets/ai-training-datasets/customer-service.textn<1K1 likes37 downloads7mo agoHugging Face28Mohaddz /Llama-Nemotron-Post-Training-Dataset-v1text100K<n<1M0 likes34 downloads2y agoHugging Face29LucidityAI /Astral-1.5-Post-Training-Dataset-SFT Astral 1.5 Post-Training Dataset A albeit smaller, yet higher-quality reasoning dataset combining mathematics, code, and general stem used in the training of the Astral 1.5 model family. Dataset Description This dataset merges four datasets to create a high quality 25 thousand example dataset. With the size of the dataset, we rely on the principle that quality > quantity leads to better model performance. Dataset Composition Setup General STEM:… See the full description on the dataset page: https://huggingface.co/datasets/LucidityAI/Astral-1.5-Post-Training-Dataset-SFT.text10K<n<100K0 likes29 downloads11mo agoHugging Face30kuzaai /training_dataset_8K_entext1K<n<10K0 likes28 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.