datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gpt-image-2-prompts-datasets
🖼️ GPT Image 2 Prompt Dataset
🖼️ The ultimate GPT Image 2 prompt dataset (5GB+). 15,000+ image generation prompts with full metadata and preview images. Truly open source: No login, no ads, no redirection. Just pure data for AI image creators.
This project is a massive collection of prompts used for OpenAI's GPT Image 2 model and the resulting generated images. The entire dataset exceeds 5GB and contains 15,000+ images, all structured into a comprehensive dataset.
Due to… See the full description on the dataset page: https://huggingface.co/datasets/Goku-OpenLab/gpt-image-2-prompts-datasets.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.fineweb10B-gpt2
fineweb10B-gpt2
This repo contains the GPT-2 tokens for fineweb10B, just as would be generated by https://github.com/KellerJordan/modded-nanogpt/tree/master (or llm.c).
You can download from this repo instead of re-tokenizing to save a couple hours of setup on a new machine.
alpaca-gpt4
Dataset Card for alpaca-gpt4
This dataset originates from this repository.
The alpaca-gpt4 dataset is specifically used for fine-tuning LLMs based on the instruction generated by GPT-4 using Alpaca prompts.
Dataset Details
Dataset Description
Each sample is comprised of four columns: instruction, input, output and text.
Language(s): English
License: Creative Commons NonCommercial (CC BY-NC 4.0)
Dataset Sources
The code from the original repository… See the full description on the dataset page: https://huggingface.co/datasets/flwrlabs/alpaca-gpt4.hle-gpt-oss-120b-no-python-260222
hle-gpt-oss-120b-no-python-260222
Deep research agent evaluation on rl-rag/hle_text_only (test split).
Results
Metric
Value
pass@4
47.9%
avg@4
26.6%
Trajectory accuracy
26.6% (2292/8632)
Questions
2158
Trajectories
8632 (4 per question)
Avg tool calls
14.5
Full conversations
❌
Model & Setup
Model
gpt-oss-120b
Judge
gpt-4o
Max tool calls
50
Temperature
0.7
Blocked domains
huggingface.co
Tool Usage… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/hle-gpt-oss-120b-no-python-260222.gpt-edit-simplerfineweb100B-gpt2marl-gpt-datasets
MARL-GPT Datasets
Offline expert trajectories from “MARL-GPT: Foundation Model for Multi-Agent Reinforcement Learning”.
Environments
This dataset includes trajectories from the three evaluation domains used in MARL-GPT: SMACv2 (StarCraft multi-agent combat), Google Research Football (GRF), and POGEMA (partially observable multi-agent pathfinding on grids).
Format
Trajectories are stored sequentially (no shuffling). Use the done flag to split the stream into… See the full description on the dataset page: https://huggingface.co/datasets/nortem/marl-gpt-datasets.Raw-GPT-traindata
Dataset Card for Dataset
This dataset has the fineweb dataset splitted in 1M rows in each files
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Each of the files ending with _p{number}.csv has 1M rows in it and they are in series
Dataset Sources
This dataset was created from HuggingFaceFW/fineweb
Uses
Has text generation data
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/WICKED4950/Raw-GPT-traindata.arc-agi3-codex-gpt5.5-su15
ARC-AGI-3 su15 — Agent Trajectories (codex-gpt5.5)
Gameplay trajectories from the harness×model pair codex-gpt5.5 playing the
ARC-AGI-3 game su15, part of the
ARA-as-world-model generalization experiment. The agent builds a structured world model
(an Agent-Native Research Artifact) live during play and consults it to crack levels it
cannot solve from cold exploration.
One dataset repo per harness×model×game: sibling repos
arc-agi3-<harness>-<model>-<game> hold the same game… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/arc-agi3-codex-gpt5.5-su15.fineweb-edu-gpt2gdpval-gpt5
GDPval with GPT-5 Execution Results
This dataset contains the OpenAI GDPval benchmark with comprehensive execution results from GPT-5, demonstrating AI capabilities across real-world professional tasks.
🎯 Dataset Overview
This is an enhanced version of the original OpenAI GDPval dataset with actual AI model execution results and professional deliverables.
📊 Key Statistics
Total tasks: 220
Tasks with AI deliverables: 87 (39.5%)
Professional files generated:… See the full description on the dataset page: https://huggingface.co/datasets/kevindenight/gdpval-gpt5.gdpval-gpt5
GDPval with GPT-5 Execution Results
This dataset contains the OpenAI GDPval benchmark with comprehensive execution results from GPT-5, demonstrating AI capabilities across real-world professional tasks.
🎯 Dataset Overview
This is an enhanced version of the original OpenAI GDPval dataset with actual AI model execution results and professional deliverables.
📊 Key Statistics
Total tasks: 220
Tasks with AI deliverables: 87 (39.5%)
Professional files generated:… See the full description on the dataset page: https://huggingface.co/datasets/lshx90/gdpval-gpt5.gdpval-gpt5-fork
GDPval Fork Dataset with GPT-5 Results
🏆 A comprehensive evaluation dataset featuring GPT-5 execution results on real-world professional tasks
This is an enhanced fork of the original OpenAI GDPval dataset with complete GPT-5 execution results, including actual deliverable files created by the AI model.
📊 Dataset Overview
Metric
Value
Total Tasks
220
AI-Completed Tasks
87 (39.5%)
Deliverable Files
492+ professional documents
Occupations
44
Industry… See the full description on the dataset page: https://huggingface.co/datasets/kevindenight/gdpval-gpt5-fork.fineweb-edu-100B-gpt2-token-shardsFineWeb Edu 100B dataset tokenized with GPT-2 tokenizer using the code in llm.c repo.
owt-gpt2
OpenWebText GPT-2 tokenized cache
This repository contains the prepared OpenWebText cache used for MDLM-MMD training and validation. The Parquet export preserves the original token IDs, attention masks, split membership, and row order. It does not retokenize, filter, or reshuffle the cached data.
Split
Rows
Tokens per row
train
8,730,826
1,024
validation
110,397
1,024
Loading
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/iasudakov/owt-gpt2.Skylion007-openwebtext-tokenizer-gpt2ChatDoctor-HealthCareMagic-Output-Improved-GPT4.1alpaca-gpt4
Dataset Card for "alpaca-gpt4"
This dataset contains English Instruction-Following generated by GPT-4 using Alpaca prompts for fine-tuning LLMs.
The dataset was originaly shared in this repository: https://github.com/Instruction-Tuning-with-GPT-4/GPT-4-LLM. This is just a wraper for compatibility with huggingface's datasets library.
Dataset structure
It contains 52K instruction-following data generated by GPT-4 using the same prompts as in Alpaca.
The dataset has the… See the full description on the dataset page: https://huggingface.co/datasets/vicgalle/alpaca-gpt4.GPT-Image-Edit-1.5M
GPT-Image-Edit-1.5M A Million-Scale, GPT-Generated Image Dataset
📃Arxiv | 🌐 Project Page | 💻Github
GPT-Image-Edit-1.5M is a comprehensive image editing dataset that is built upon HQ-Edit, UltraEdit, OmniEdit and Complex-Edit, with all output images regenerated with GPT-Image-1.
📣 News
[2025.08.20] 🚀 We provide a script for multi-process downloading. See Multi-process Download.
[2025.07.27] 🤗 We release GPT-Image-Edit, a state-of-the-art image editing model with… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/GPT-Image-Edit-1.5M.Taur_CoT_Analysis_Project___gpt-4o-2024-08-06arc-agi3-codex-gpt5.6sol-ls20
ARC-AGI-3 ls20 — Agent Trajectories (codex-gpt5.6sol)
Gameplay trajectories from the harness×model pair codex-gpt5.6sol playing the
ARC-AGI-3 game ls20, part of the
ARA-as-world-model generalization experiment. The agent builds a structured world model
(an Agent-Native Research Artifact) live during play and consults it to crack levels it
cannot solve from cold exploration.
One dataset repo per harness×model×game: sibling repos
arc-agi3-<harness>-<model>-<game> hold the same… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/arc-agi3-codex-gpt5.6sol-ls20.vibevoice-gptinformal_persian-single-speakergpt-oss-20b-moe-expert-power-traces-320k
GPT-OSS-20B MoE Expert Power Traces (320k, ChipWhisperer)
This dataset contains analog power traces captured with a ChipWhisperer Husky while running forced single-expert MoE computations derived from openai/gpt-oss-20b on an NVIDIA H100.
What is recorded
Each trace corresponds to one capture trial where:
A fixed expert id is selected (expert_00 ... expert_31).
A random hidden-state tensor is generated once per trial.
The selected expert computation is executed… See the full description on the dataset page: https://huggingface.co/datasets/masterpieceexternal/gpt-oss-20b-moe-expert-power-traces-320k.GPT-NL_Public_Corpus
Dataset Card GPT-NL Public Corpus
The GPT-NL Public Corpus is the largest permissively licensed Dutch-language resource available for large language model pretraining. It consists of 29 curated collections totaling over 524 billion tokens, including 36B Dutch, 207B English, 232B code, and 48B German/Danish tokens. All data is sourced under permissive licensing and redistributed under a CC-BY license.
For more details please refer to our Public Corpus article.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/GPT-NL/GPT-NL_Public_Corpus.2026-09-28-odcv-gptoss120b-0-nosynth
odcv eval of tinker://cdd100d5-e08c-52a7-9ba9-890187111a57:train:0/sampler_weights/2026-09-28-gptoss120b-0-nosynth (mode=harmony_medium)
field
value
experiment
odcv eval of tinker://cdd100d5-e08c-52a7-9ba9-890187111a57:train:0/sampler_weights/2026-09-28-gptoss120b-0-nosynth (mode=harmony_medium)
date_generated
2026-09-28
constitution
none
source_repo
teaching_claude_why_replication @ 779625e8cd54d56017dac2141bd23e6ff407022d
models
{"target":… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-28-odcv-gptoss120b-0-nosynth.processed_gpt_dataset_big
Dataset Card for "processed_gpt_dataset_big"
More Information needed
arc-agi3-codex-gpt5.5-s5i5
ARC-AGI-3 s5i5 — Agent Trajectories (codex-gpt5.5)
Gameplay trajectories from the harness×model pair codex-gpt5.5 playing the
ARC-AGI-3 game s5i5, part of the
ARA-as-world-model generalization experiment. The agent builds a structured world model
(an Agent-Native Research Artifact) live during play and consults it to crack levels it
cannot solve from cold exploration.
One dataset repo per harness×model×game: sibling repos
arc-agi3-<harness>-<model>-<game> hold the same game… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/arc-agi3-codex-gpt5.5-s5i5.2wikimultihopqa_with_q_gpt35
2WikiMultihopQA Dataset with GPT-3.5 Generated Questions
Overview
This repository hosts an enhanced version of the 2WikiMultihopQA dataset, where each supporting sentence in the dataset has been supplemented with questions generated using OpenAI's GPT-3.5 turbo API. The aim is to provide a richer context for each entry, potentially benefiting various NLP tasks, such as question answering and context understanding.
Dataset Format
Each entry in the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/scholarly-shadows-syndicate/2wikimultihopqa_with_q_gpt35.copyright_gpt_neo_1_3B
