Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Goku-OpenLab /gpt-image-2-prompts-datasets 🖼️ GPT Image 2 Prompt Dataset 🖼️ The ultimate GPT Image 2 prompt dataset (5GB+). 15,000+ image generation prompts with full metadata and preview images. Truly open source: No login, no ads, no redirection. Just pure data for AI image creators. This project is a massive collection of prompts used for OpenAI's GPT Image 2 model and the resulting generated images. The entire dataset exceeds 5GB and contains 15,000+ images, all structured into a comprehensive dataset. Due to… See the full description on the dataset page: https://huggingface.co/datasets/Goku-OpenLab/gpt-image-2-prompts-datasets.imagetext-to-image10K<n<100K8 likes111k downloads4d agoHugging Face02Manusagents /GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset 📖 The Open Distillation Codex 🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌 Where 73 open-source minds converge into one unified stream of intelligence 18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+ "We did not write this dataset. We assembled it. Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing. Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.texttext-generation10M<n<100M215 likes18k downloads1mo agoHugging Face03kjj0 /fineweb10B-gpt2 fineweb10B-gpt2 This repo contains the GPT-2 tokens for fineweb10B, just as would be generated by https://github.com/KellerJordan/modded-nanogpt/tree/master (or llm.c). You can download from this repo instead of re-tokenizing to save a couple hours of setup on a new machine. 12 likes13k downloads2y agoHugging Face04flwrlabs /alpaca-gpt4 Dataset Card for alpaca-gpt4 This dataset originates from this repository. The alpaca-gpt4 dataset is specifically used for fine-tuning LLMs based on the instruction generated by GPT-4 using Alpaca prompts. Dataset Details Dataset Description Each sample is comprised of four columns: instruction, input, output and text. Language(s): English License: Creative Commons NonCommercial (CC BY-NC 4.0) Dataset Sources The code from the original repository… See the full description on the dataset page: https://huggingface.co/datasets/flwrlabs/alpaca-gpt4.texttext-generation10K<n<100K0 likes12k downloads2y agoHugging Face05rl-rag /hle-gpt-oss-120b-no-python-260222 hle-gpt-oss-120b-no-python-260222 Deep research agent evaluation on rl-rag/hle_text_only (test split). Results Metric Value pass@4 47.9% avg@4 26.6% Trajectory accuracy 26.6% (2292/8632) Questions 2158 Trajectories 8632 (4 per question) Avg tool calls 14.5 Full conversations ❌ Model & Setup Model gpt-oss-120b Judge gpt-4o Max tool calls 50 Temperature 0.7 Blocked domains huggingface.co Tool Usage… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/hle-gpt-oss-120b-no-python-260222.tabular1K<n<10K1 likes11k downloads8mo agoHugging Face06UCSC-VLAA /gpt-edit-simplerimage1M<n<10M14 likes10k downloads1y agoHugging Face07kjj0 /fineweb100B-gpt21 likes9.1k downloads2y agoHugging Face08nortem /marl-gpt-datasets MARL-GPT Datasets Offline expert trajectories from “MARL-GPT: Foundation Model for Multi-Agent Reinforcement Learning”. Environments This dataset includes trajectories from the three evaluation domains used in MARL-GPT: SMACv2 (StarCraft multi-agent combat), Google Research Football (GRF), and POGEMA (partially observable multi-agent pathfinding on grids). Format Trajectories are stored sequentially (no shuffling). Use the done flag to split the stream into… See the full description on the dataset page: https://huggingface.co/datasets/nortem/marl-gpt-datasets.tabularreinforcement-learning100M<n<1B0 likes7.1k downloads7mo agoHugging Face09WICKED4950 /Raw-GPT-traindata Dataset Card for Dataset This dataset has the fineweb dataset splitted in 1M rows in each files This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Each of the files ending with _p{number}.csv has 1M rows in it and they are in series Dataset Sources This dataset was created from HuggingFaceFW/fineweb Uses Has text generation data Dataset… See the full description on the dataset page: https://huggingface.co/datasets/WICKED4950/Raw-GPT-traindata.text-generation10M<n<100M3 likes6.4k downloads5mo agoHugging Face10AgentNativeResearchLab /arc-agi3-codex-gpt5.5-su15 ARC-AGI-3 su15 — Agent Trajectories (codex-gpt5.5) Gameplay trajectories from the harness×model pair codex-gpt5.5 playing the ARC-AGI-3 game su15, part of the ARA-as-world-model generalization experiment. The agent builds a structured world model (an Agent-Native Research Artifact) live during play and consults it to crack levels it cannot solve from cold exploration. One dataset repo per harness×model×game: sibling repos arc-agi3-<harness>-<model>-<game> hold the same game… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/arc-agi3-codex-gpt5.5-su15.reinforcement-learning0 likes6.3k downloads1mo agoHugging Face11kushalt /fineweb-edu-gpt2tabular10M<n<100M0 likes5.7k downloads8mo agoHugging Face12kevindenight /gdpval-gpt5 GDPval with GPT-5 Execution Results This dataset contains the OpenAI GDPval benchmark with comprehensive execution results from GPT-5, demonstrating AI capabilities across real-world professional tasks. 🎯 Dataset Overview This is an enhanced version of the original OpenAI GDPval dataset with actual AI model execution results and professional deliverables. 📊 Key Statistics Total tasks: 220 Tasks with AI deliverables: 87 (39.5%) Professional files generated:… See the full description on the dataset page: https://huggingface.co/datasets/kevindenight/gdpval-gpt5.documentothern<1K0 likes5.1k downloads11mo agoHugging Face13lshx90 /gdpval-gpt5 GDPval with GPT-5 Execution Results This dataset contains the OpenAI GDPval benchmark with comprehensive execution results from GPT-5, demonstrating AI capabilities across real-world professional tasks. 🎯 Dataset Overview This is an enhanced version of the original OpenAI GDPval dataset with actual AI model execution results and professional deliverables. 📊 Key Statistics Total tasks: 220 Tasks with AI deliverables: 87 (39.5%) Professional files generated:… See the full description on the dataset page: https://huggingface.co/datasets/lshx90/gdpval-gpt5.documentothern<1K0 likes4.9k downloads6mo agoHugging Face14kevindenight /gdpval-gpt5-fork GDPval Fork Dataset with GPT-5 Results 🏆 A comprehensive evaluation dataset featuring GPT-5 execution results on real-world professional tasks This is an enhanced fork of the original OpenAI GDPval dataset with complete GPT-5 execution results, including actual deliverable files created by the AI model. 📊 Dataset Overview Metric Value Total Tasks 220 AI-Completed Tasks 87 (39.5%) Deliverable Files 492+ professional documents Occupations 44 Industry… See the full description on the dataset page: https://huggingface.co/datasets/kevindenight/gdpval-gpt5-fork.documentothern<1K0 likes4.7k downloads1y agoHugging Face15karpathy /fineweb-edu-100B-gpt2-token-shardsFineWeb Edu 100B dataset tokenized with GPT-2 tokenizer using the code in llm.c repo. 8 likes4.6k downloads2y agoHugging Face16iasudakov /owt-gpt2 OpenWebText GPT-2 tokenized cache This repository contains the prepared OpenWebText cache used for MDLM-MMD training and validation. The Parquet export preserves the original token IDs, attention masks, split membership, and row order. It does not retokenize, filter, or reshuffle the cached data. Split Rows Tokens per row train 8,730,826 1,024 validation 110,397 1,024 Loading from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/iasudakov/owt-gpt2.timeseries1M<n<10M0 likes3.9k downloads8d agoHugging Face17apollo-research /Skylion007-openwebtext-tokenizer-gpt21M<n<10M3 likes3.8k downloads3y agoHugging Face18hoanganhpham /ChatDoctor-HealthCareMagic-Output-Improved-GPT4.1text10K<n<100K1 likes3.6k downloads1y agoHugging Face19vicgalle /alpaca-gpt4 Dataset Card for "alpaca-gpt4" This dataset contains English Instruction-Following generated by GPT-4 using Alpaca prompts for fine-tuning LLMs. The dataset was originaly shared in this repository: https://github.com/Instruction-Tuning-with-GPT-4/GPT-4-LLM. This is just a wraper for compatibility with huggingface's datasets library. Dataset structure It contains 52K instruction-following data generated by GPT-4 using the same prompts as in Alpaca. The dataset has the… See the full description on the dataset page: https://huggingface.co/datasets/vicgalle/alpaca-gpt4.texttext-generation10K<n<100K326 likes3.4k downloads3y agoHugging Face20UCSC-VLAA /GPT-Image-Edit-1.5M GPT-Image-Edit-1.5M A Million-Scale, GPT-Generated Image Dataset 📃Arxiv | 🌐 Project Page | 💻Github GPT-Image-Edit-1.5M is a comprehensive image editing dataset that is built upon HQ-Edit, UltraEdit, OmniEdit and Complex-Edit, with all output images regenerated with GPT-Image-1. 📣 News [2025.08.20] 🚀 We provide a script for multi-process downloading. See Multi-process Download. [2025.07.27] 🤗 We release GPT-Image-Edit, a state-of-the-art image editing model with… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/GPT-Image-Edit-1.5M.imageimage-to-image1M<n<10M90 likes3.4k downloads1y agoHugging Face21TAUR-Lab /Taur_CoT_Analysis_Project___gpt-4o-2024-08-06text10K<n<100K1 likes3.4k downloads2y agoHugging Face22AgentNativeResearchLab /arc-agi3-codex-gpt5.6sol-ls20 ARC-AGI-3 ls20 — Agent Trajectories (codex-gpt5.6sol) Gameplay trajectories from the harness×model pair codex-gpt5.6sol playing the ARC-AGI-3 game ls20, part of the ARA-as-world-model generalization experiment. The agent builds a structured world model (an Agent-Native Research Artifact) live during play and consults it to crack levels it cannot solve from cold exploration. One dataset repo per harness×model×game: sibling repos arc-agi3-<harness>-<model>-<game> hold the same… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/arc-agi3-codex-gpt5.6sol-ls20.textreinforcement-learning0 likes3.4k downloads1mo agoHugging Face23Dorsaasgari /vibevoice-gptinformal_persian-single-speakeraudio1K<n<10K0 likes3.3k downloads27d agoHugging Face24masterpieceexternal /gpt-oss-20b-moe-expert-power-traces-320k GPT-OSS-20B MoE Expert Power Traces (320k, ChipWhisperer) This dataset contains analog power traces captured with a ChipWhisperer Husky while running forced single-expert MoE computations derived from openai/gpt-oss-20b on an NVIDIA H100. What is recorded Each trace corresponds to one capture trial where: A fixed expert id is selected (expert_00 ... expert_31). A random hidden-state tensor is generated once per trial. The selected expert computation is executed… See the full description on the dataset page: https://huggingface.co/datasets/masterpieceexternal/gpt-oss-20b-moe-expert-power-traces-320k.audio-classification100K<n<1M0 likes3.3k downloads4mo agoHugging Face25GPT-NL /GPT-NL_Public_Corpus Dataset Card GPT-NL Public Corpus The GPT-NL Public Corpus is the largest permissively licensed Dutch-language resource available for large language model pretraining. It consists of 29 curated collections totaling over 524 billion tokens, including 36B Dutch, 207B English, 232B code, and 48B German/Danish tokens. All data is sourced under permissive licensing and redistributed under a CC-BY license. For more details please refer to our Public Corpus article. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/GPT-NL/GPT-NL_Public_Corpus.tabular100M<n<1B22 likes3.2k downloads1mo agoHugging Face26dougalldeepmind /2026-09-28-odcv-gptoss120b-0-nosynth odcv eval of tinker://cdd100d5-e08c-52a7-9ba9-890187111a57:train:0/sampler_weights/2026-09-28-gptoss120b-0-nosynth (mode=harmony_medium) field value experiment odcv eval of tinker://cdd100d5-e08c-52a7-9ba9-890187111a57:train:0/sampler_weights/2026-09-28-gptoss120b-0-nosynth (mode=harmony_medium) date_generated 2026-09-28 constitution none source_repo teaching_claude_why_replication @ 779625e8cd54d56017dac2141bd23e6ff407022d models {"target":… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-28-odcv-gptoss120b-0-nosynth.text0 likes2.9k downloads11d agoHugging Face27sanagnos /processed_gpt_dataset_big Dataset Card for "processed_gpt_dataset_big" More Information needed 1M<n<10M0 likes2.9k downloads4y agoHugging Face28AgentNativeResearchLab /arc-agi3-codex-gpt5.5-s5i5 ARC-AGI-3 s5i5 — Agent Trajectories (codex-gpt5.5) Gameplay trajectories from the harness×model pair codex-gpt5.5 playing the ARC-AGI-3 game s5i5, part of the ARA-as-world-model generalization experiment. The agent builds a structured world model (an Agent-Native Research Artifact) live during play and consults it to crack levels it cannot solve from cold exploration. One dataset repo per harness×model×game: sibling repos arc-agi3-<harness>-<model>-<game> hold the same game… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/arc-agi3-codex-gpt5.5-s5i5.reinforcement-learning0 likes2.9k downloads1mo agoHugging Face29scholarly-shadows-syndicate /2wikimultihopqa_with_q_gpt35 2WikiMultihopQA Dataset with GPT-3.5 Generated Questions Overview This repository hosts an enhanced version of the 2WikiMultihopQA dataset, where each supporting sentence in the dataset has been supplemented with questions generated using OpenAI's GPT-3.5 turbo API. The aim is to provide a richer context for each entry, potentially benefiting various NLP tasks, such as question answering and context understanding. Dataset Format Each entry in the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/scholarly-shadows-syndicate/2wikimultihopqa_with_q_gpt35.text10K<n<100K2 likes2.8k downloads3y agoHugging Face30bxiong /copyright_gpt_neo_1_3B0 likes2.7k downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.