Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jat-project /jat-dataset JAT Dataset Dataset Description The Jack of All Trades (JAT) dataset combines a wide range of individual datasets. It includes expert demonstrations by expert RL agents, image and caption pairs, textual data and more. The JAT dataset is part of the JAT project, which aims to build a multimodal generalist agent. Paper: https://huggingface.co/papers/2402.09844 Usage >>> from datasets import load_dataset >>> dataset =… See the full description on the dataset page: https://huggingface.co/datasets/jat-project/jat-dataset.imagereinforcement-learning100M<n<1B78 likes202k downloads3y agoHugging Face02llamafactory /demo_data 1,000 examples from https://huggingface.co/datasets/llamafactory/alpaca_gpt4_en 1,000 examples from https://huggingface.co/datasets/llamafactory/alpaca_gpt4_zh 300 examples from https://huggingface.co/datasets/llamafactory/glaive_toolcall_en 300 examples from https://huggingface.co/datasets/llamafactory/glaive_toolcall_zh 91 examples for identity learning 300 examples from https://huggingface.co/datasets/cognitivecomputations/SystemChat-2.0 6 examples for multimodal supervised… See the full description on the dataset page: https://huggingface.co/datasets/llamafactory/demo_data.texttext-generation1K<n<10K1 likes99k downloads2y agoHugging Face03llamafactory /tiny-supervised-datasettexttext-generationn<1K4 likes46k downloads2y agoHugging Face04zgcagi /ZGCM-1-Datagated A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search Zhongguancun Academy · Zhongguancun Institute of Artificial Intelligence 📄 Tech Report · 🤗 Model · 🤗 Data · 📈 Training Log 📊 Results · 💻 Training Code · 💬 WeChat Community Introduction ZGCM-1 is a 7.39B-parameter dense language model trained from scratch, built for mathematical reasoning and tool-assisted search. It combines deliberate internal thinking with active information… See the full description on the dataset page: https://huggingface.co/datasets/zgcagi/ZGCM-1-Data.texttext-generation1B<n<10B88 likes45k downloads14d agoHugging Face05DataMuncher-Labs /UltiMath Dataset Card for UltiMath UltiMath is a large-scale synthetic dataset containing ~33 billion math reasoning examples, designed to enhance arithmetic and symbolic reasoning in large language models (LLMs). Dataset Details Dataset Description Curated by: [Roman] Funded by: [No funding used] Shared by [Roman]: [Uploads via API] License: [CC by SA 4.0] Dataset Sources [Code Generated] Uses Designed to improve multi-step arithmetic… See the full description on the dataset page: https://huggingface.co/datasets/DataMuncher-Labs/UltiMath.texttext-generation10B<n<100B46 likes33k downloads9mo agoHugging Face06OpenLLM-France /Lucie-Training-Dataset Lucie Training Dataset Card The Lucie Training Dataset is a curated collection of text data in English, French, German, Spanish and Italian culled from a variety of sources including: web data, video subtitles, academic papers, digital books, newspapers, and magazines, some of which were processed by Optical Character Recognition (OCR). It also contains samples of diverse programming languages. The Lucie Training Dataset was used to pretrain Lucie-7B, a foundation LLM with… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Lucie-Training-Dataset.texttext-generation10B<n<100B39 likes31k downloads1y agoHugging Face07Chrisneverdie /OnlySports_Dataset 🏀nlySports Dataset Overview OnlySports Dataset is a comprehensive collection of English sports documents, comprising a diverse range of content including news articles, blogs, match reports, interviews, and tutorials. This dataset is part of the larger OnlySports collection, which includes: OnlySportsLM: A 196M parameter sports-domain language model OnlySports Dataset: The dataset described in this README OnlySports Benchmark: A novel evaluation method for assessing… See the full description on the dataset page: https://huggingface.co/datasets/Chrisneverdie/OnlySports_Dataset.texttext-generation1B<n<10B5 likes29k downloads2y agoHugging Face08tensorshield /reddit_dataset_157 Bittensor Subnet 13 Reddit Dataset Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed Reddit data. The data is continuously updated by network miners, providing a real-time stream of Reddit content for various analytical and machine learning tasks. For more information about the dataset, please visit the official repository. Supported Tasks The versatility of this dataset allows… See the full description on the dataset page: https://huggingface.co/datasets/tensorshield/reddit_dataset_157.texttext-classification10M<n<100M4 likes26k downloads2y agoHugging Face09open-law-data-thailand /soc-ratchakitcha Royal Gazette Thailand (Ratchakitcha) Dataset ชุดข้อมูลราชกิจจานุเบกษา (แบบ Machine Readable) โครงการ Open Law Data Thailand ร่วมกับคณะกรรมาธิการการพาณิชย์และการอุตสาหกรรม วุฒิสภา ได้รับความอนุเคราะห์ข้อมูลจาก สำนักเลขาธิการคณะรัฐมนตรี (สลค.) เพื่อเผยแพร่ข้อมูลกฎหมายไทยสู่สาธารณะในรูปแบบที่ประมวลผลได้ด้วยคอมพิวเตอร์ (Machine Readable) เพื่อส่งเสริมนวัตกรรม Legal Tech และ AI ของประเทศไทย Dataset Description ชุดข้อมูลนี้รวบรวมรายการประกาศในราชกิจจานุเบกษา… See the full description on the dataset page: https://huggingface.co/datasets/open-law-data-thailand/soc-ratchakitcha.tabulartext-retrieval1M<n<10M14 likes24k downloads4h agoHugging Face10chewwt /po_qwen14b_tabular_data BoLT Prompt Optimization — Tabular Dataset For prompt optimization tasks in BoLT, an accessible benchmark for black-box optimization on LLM tasks. Dataset Description The dataset covers 5,014 evaluated instructions. Each row is a candidate system-prompt instruction paired with its empirically measured MATH-500 (4-shot, non-thinking mode) scores. Evaluation details: Model: Qwen/Qwen3-14B Task: minerva_math500 (4-shot) (from lm-eval library) System prompt:… See the full description on the dataset page: https://huggingface.co/datasets/chewwt/po_qwen14b_tabular_data.tabulartext-generation1K<n<10K1 likes21k downloads5mo agoHugging Face11Manusagents /GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset 📖 The Open Distillation Codex 🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌 Where 73 open-source minds converge into one unified stream of intelligence 18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+ "We did not write this dataset. We assembled it. Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing. Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.texttext-generation10M<n<100M216 likes18k downloads1mo agoHugging Face12AyoubChLin /Company-document-dataset-v2 Company Documents v2 Generation complete: all 13 document types have completed export and upload checkpoints. Synthetic, born-digital business documents rendered from four open sample databases, with exact gold labels: 353,580 PDFs (404,514 pages) of 13 document types in English and French, issued by 60 synthetic companies, each with its own letterhead, numbering and wording. Successor of CompanyDocuments (2,677 PDFs, 4 types). Dataset overview property… See the full description on the dataset page: https://huggingface.co/datasets/AyoubChLin/Company-document-dataset-v2.documentdocument-question-answering100K<n<1M2 likes12k downloads3d agoHugging Face13mohameddalii /coda-llm-data Coda LLM Project & Dataset Repository This repository contains the full end-to-end dataset, fine-tuning scripts, evaluation suites, load testing harness, and proxy architecture for Coda LLM (Granite-4.2-8B Najdi Sales Agent). Model Repository: mohameddalii/coda-llm Dataset / Code Repository: mohameddalii/coda-llm-data 📁 Repository Structure coda-llm-data/ ├── data/ │ ├── raw/ # Raw generated multi-turn dialogues across domains │ ├──… See the full description on the dataset page: https://huggingface.co/datasets/mohameddalii/coda-llm-data.texttext-generation1K<n<10K0 likes12k downloads41m agoHugging Face14databricks /officeqagated OfficeQA Dataset Summary OfficeQA is a grounded reasoning benchmark by Databricks for evaluating model and agent performance on end-to-end reasoning over real-world documents. The benchmark consists of question–answer pairs that require reasoning over historical U.S. Treasury Bulletin documents (1939–2025), which contain dense financial tables, charts, and narrative text. OfficeQA is designed to test retrieval, tool use, and multi-step reasoning in… See the full description on the dataset page: https://huggingface.co/datasets/databricks/officeqa.documentquestion-answeringn<1K30 likes9k downloads3mo agoHugging Face15arnizamani /Sindhi-texts-big-dataset Sindhi Texts (big dataset) A large plain-text corpus of Sindhi (سنڌي), assembled for pretraining language models. It combines material digitized by Sindhi literary institutions and forums, a Sindhi encyclopedia, newspaper archives, a classical dictionary, and the Sindhi portions of two web-crawl corpora. 3.19 GB, ~1.81 billion characters, ~390,000 documents across 9 sources. With a Sindhi-specific 12k SentencePiece tokenizer that is roughly 530M tokens (3.2–3.5 characters per… See the full description on the dataset page: https://huggingface.co/datasets/arnizamani/Sindhi-texts-big-dataset.texttext-generation100K<n<1M4 likes8k downloads2mo agoHugging Face16tdr-research-team /tdr-dataset TDR: Tagalog Diacritic Restoration Dataset TDR is a manually annotated corpus of contemporary Tagalog news text covering 40 frequent homographs, with diacritical marks (tuldik) restored to disambiguate meaning and pronunciation. Dataset Description Tagalog uses diacritics (tuldik) to mark syllable stress and final glottal stops, resolving ambiguity among homographs — words that share the same spelling but differ in meaning and pronunciation. For example, puno can… See the full description on the dataset page: https://huggingface.co/datasets/tdr-research-team/tdr-dataset.texttext-classification10K<n<100K0 likes7.2k downloads4d agoHugging Face17takschdube /moltbook-dataset Moltbook Dataset A longitudinal dataset of social interactions from Moltbook — an AI-agent social platform where autonomous "Molties" post, comment, and interact. Collected automatically and published as timestamped snapshots for temporal analysis. Dataset Statistics Metric Count Posts (platform total) 4,427,314 Comments (platform total) 13,538,160 Posts (collected) 422,497 Comments (collected) 3,723,066 Agents 57,217 Social graph edges 826… See the full description on the dataset page: https://huggingface.co/datasets/takschdube/moltbook-dataset.tabulartext-generation1M<n<10M4 likes6.5k downloads2h agoHugging Face18argilla /ifeval-like-data IFEval Like Data This dataset contains instruction-response pairs synthetically generated using Qwen/Qwen2.5-72B-Instruct following the style of google/IFEval dataset and verified for correctness with lm-evaluation-harness. The dataset contains two subsets: default: which contains 550k unfiltered rows synthetically generated with Qwen2.5-72B-Instruct, a few system prompts and MagPie prompting technique. The prompts can contain conflicting instructions as defined in… See the full description on the dataset page: https://huggingface.co/datasets/argilla/ifeval-like-data.texttext-generation100K<n<1M50 likes6.2k downloads2y agoHugging Face19lesserfield /4chan-datasetsPlease see repo to turn the text file into json/csv format Deleted some boards, since they are already archived by https://archive.4plebs.org/ texttext-generation34 likes5.6k downloads3y agoHugging Face20lingshu-medical-mllm /lingshu_training_data_medical_domain Website &nbsp;&nbsp; 🤖 7B Model &nbsp;&nbsp; 🤖 8B Model based on InternVL3 &nbsp;&nbsp; 🤖 32B Model &nbsp;&nbsp; MedEvalKit &nbsp;&nbsp; Technical Report &nbsp;&nbsp; Lingshu MCP Lingshu Medical MLLM Training Data (Medical Domain) This dataset contains the medical-domain training data used in the multi-stage training of the Lingshu Medical Multimodal Large Language Model (MLLM). General-domain data has been removed; only medical data is included. The training… See the full description on the dataset page: https://huggingface.co/datasets/lingshu-medical-mllm/lingshu_training_data_medical_domain.textimage-to-text100M<n<1B8 likes5.4k downloads1mo agoHugging Face21SHSLab /Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection 🧬 Omni-Frontier Collection Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible. 📖 Jump to What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/SHSLab/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.tabulartext-generation10M<n<100M5 likes5.3k downloads1mo agoHugging Face22johanneskirmayr /car-bench-dataset CAR-Bench Dataset CAR-Bench is a benchmark for evaluating AI voice assistants in a realistic automotive (car) environment. It tests an agent's ability to correctly use vehicle control tools, handle disambiguation, and avoid hallucinations. Dataset Structure The dataset is organized into task configs and mock data configs: Tasks Each task defines a user persona, an instruction, the initial vehicle/environment context, and the ground-truth sequence of tool-call… See the full description on the dataset page: https://huggingface.co/datasets/johanneskirmayr/car-bench-dataset.tabulartext-generation1M<n<10M4 likes5k downloads8mo agoHugging Face23Trendyol /Trendyol-Cybersecurity-Instruction-Tuning-Dataset Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0) 🚀 TL;DR 53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/Trendyol/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.texttext-generation10K<n<100K135 likes5k downloads1y agoHugging Face24Gunulhona /llm_datasetstexttext-generation100K<n<1M0 likes4.8k downloads3y agoHugging Face25futuremoon /x_dataset_39 Bittensor Subnet 13 X (Twitter) Dataset Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks. For more information about the dataset, please visit the official repository. Supported Tasks The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/futuremoon/x_dataset_39.texttext-classification1B<n<10B2 likes4.3k downloads1y agoHugging Face26PrimeIntellect /SWE-Lego-Real-Data-Verified SWE-Lego-Real-Data-Verified Gold-patch-validated subset of PrimeIntellect/SWE-Lego-Real-Data (itself a fixed fork of SWE-Lego's real-data split). The resolved split contains 4,323 / 4,432 rows (97.54%) verified scoreable end-to-end: apply test_patch, apply the gold patch, run the row's test_cmd in its image, require every F2P/P2P test to report PASSED. Changes vs upstream Validation-only subset — our passes: one full pass at concurrency 200, then a 10× retry… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/SWE-Lego-Real-Data-Verified.texttext-generation1K<n<10K1 likes3.9k downloads8d agoHugging Face27AlicanKiraz0 /Cybersecurity-Dataset-Fenrir-v2.1 Cybersecurity Defense Instruction-Tuning Dataset (v2.1) Created by Alican Kiraz TL;DR A ready-to-train dataset of 99,870 high-quality system / user / assistant triples for defensive, alignment-safe cybersecurity SFT training. Apache-2.0 licensed and production-ready. Scope: OWASP Top 10, MITRE ATT&CK, NIST CSF, CIS Controls, ASD Essential 8, modern authentication (OAuth 2 / OIDC / SAML), SSL / TLS, Cloud & DevSecOps, Cryptography, and AI Security. 1  What’s… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Cybersecurity-Dataset-Fenrir-v2.1.texttext-generation10K<n<100K153 likes3.6k downloads6mo agoHugging Face28Weoxx62 /lauma-dataset 🚀 Lauma Dataset (51.6 GB) Bienvenue sur le dépôt officiel du Lauma Dataset. Ce jeu de données massif de 51.6 Go est une mine d'or hautement optimisée pour le pré-entraînement, le raffinement (Fine-Tuning) et l'alignement de modèles de langage (LLMs) francophones et hybrides. Il regroupe un mélange massif de textes scientifiques, de code source, de chaînes de raisonnement profond (Reasoning/Thinking), de dialogues avancés et de culture générale. 🛠️ Spécifications… See the full description on the dataset page: https://huggingface.co/datasets/Weoxx62/lauma-dataset.texttext-generation1K<n<10K0 likes3.5k downloads4mo agoHugging Face29StormKing99 /x_dataset_8191 Bittensor Subnet 13 X (Twitter) Dataset Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks. For more information about the dataset, please visit the official repository. Supported Tasks The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/StormKing99/x_dataset_8191.texttext-classification100M<n<1B0 likes3.4k downloads1y agoHugging Face30croissantllm /croissant_dataset CroissantLLM: A Truly Bilingual French-English Language Model Dataset https://arxiv.org/abs/2402.00786 Licenses Data redistributed here is subject to the original license under which it was collected. All license information is detailed in the Data section of the Technical report. Citation @misc{faysse2024croissantllm, title={CroissantLLM: A Truly Bilingual French-English Language Model}, author={Manuel Faysse and Patrick Fernandes and… See the full description on the dataset page: https://huggingface.co/datasets/croissantllm/croissant_dataset.texttranslation10B<n<100B8 likes3.3k downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.