datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GPT-5.5-CodexThis dataset was generated using teich by TeichAI
GPT-5.5 Agent traces
This directory contains raw agent trace files generated by teich.
JSONL files: 317
Model metadata: gpt-5.5
Training-ready tools
Generated agent traces carry configured or recovered tool schemas so tools remain available for training even when a session did not call them.
Native Claude Code imports recover schemas for Claude Code and Claude Desktop built-ins, plus conservative name-derived MCP… See the full description on the dataset page: https://huggingface.co/datasets/AletheiaResearch/GPT-5.5-Codex.share-codex
Local Coding Agent Sessions
Local coding-agent session transcripts exported with
share-codex.
Rows are in train.jsonl. Each row contains id, ordered messages, and
metadata. Messages include user prompts, assistant responses/tool calls, and
tool outputs. Internal prompts and reasoning are excluded by default.
Dataset Statistics
Statistics are computed from the final merged train.jsonl rows at export time.
Metric
Value
Sessions
4019
User turns
16639… See the full description on the dataset page: https://huggingface.co/datasets/nmuendler/share-codex.Kimi-K3-CodexThis dataset was generated using teich by TeichAI
Kimi-K3 Codex traces
This directory contains raw agent trace files generated by teich.
JSONL files: 5
Model metadata: moonshotai/kimi-k3
Training-ready tools
Generated agent traces carry configured or recovered tool schemas so tools remain available for training even when a session did not call them.
Native Claude Code imports recover schemas for Claude Code and Claude Desktop built-ins, plus conservative… See the full description on the dataset page: https://huggingface.co/datasets/AletheiaResearch/Kimi-K3-Codex.my-personal-codex-data
Coding Agent Conversation Logs
This is a performance art project. Anthropic built their models on the world's freely shared information, then introduced increasingly dystopian data policies to stop anyone else from doing the same with their data — pulling up the ladder behind them. DataClaw lets you throw the ladder back down. The dataset it produces is yours to share.
Exported with DataClaw.
Tag: dataclaw — Browse all DataClaw datasets
Stats
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/xuechengjiang/my-personal-codex-data.my-personal-codex-data
Coding Agent Conversation Logs
This is a performance art project. Anthropic built their models on the world's freely shared information, then introduced increasingly dystopian data policies to stop anyone else from doing the same with their data — pulling up the ladder behind them. DataClaw lets you throw the ladder back down. The dataset it produces is yours to share.
Exported with DataClaw.
Tag: dataclaw — Browse all DataClaw datasets
Stats
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/michaelwaves/my-personal-codex-data.my-personal-codex-data
Coding Agent Conversation Logs
This is a performance art project. Anthropic built their models on the world's freely shared information, then introduced increasingly dystopian data policies to stop anyone else from doing the same with their data - pulling up the ladder behind them. DataClaw lets you throw the ladder back down. The dataset it produces is yours to share.
Exported with DataClaw.
Tag: dataclaw - Browse all DataClaw datasets
Stats
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/wop/my-personal-codex-data.my-personal-codex-data
Coding Agent Conversation Logs
This is a performance art project. Anthropic built their models on the world's freely shared information, then introduced increasingly dystopian data policies to stop anyone else from doing the same with their data — pulling up the ladder behind them. DataClaw lets you throw the ladder back down. The dataset it produces is yours to share.
Exported with DataClaw.
Tag: dataclaw — Browse all DataClaw datasets
Stats
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/akenove/my-personal-codex-data.my-personal-codex-data
Coding Agent Conversation Logs
This is a performance art project. Anthropic built their models on the world's freely shared information, then introduced increasingly dystopian data policies to stop anyone else from doing the same with their data - pulling up the ladder behind them. DataClaw lets you throw the ladder back down. The dataset it produces is yours to share.
Exported with DataClaw.
Tag: dataclaw - Browse all DataClaw datasets
Stats
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/masterda/my-personal-codex-data.my-personal-codex-data
Coding Agent Conversation Logs
This is a performance art project. Anthropic built their models on the world's freely shared information, then introduced increasingly dystopian data policies to stop anyone else from doing the same with their data - pulling up the ladder behind them. DataClaw lets you throw the ladder back down. The dataset it produces is yours to share.
Exported with DataClaw.
Tag: dataclaw - Browse all DataClaw datasets
Stats
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/sinhaankur/my-personal-codex-data.my-personal-codex-data
Coding Agent Conversation Logs
This is a performance art project. Anthropic built their models on the world's freely shared information, then introduced increasingly dystopian data policies to stop anyone else from doing the same with their data - pulling up the ladder behind them. DataClaw lets you throw the ladder back down. The dataset it produces is yours to share.
Exported with DataClaw.
Tag: dataclaw - Browse all DataClaw datasets
Stats
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/wuuski/my-personal-codex-data.hackathon-advisor-codex-traces
Hackathon Advisor Codex Session Traces
Real Codex session logs for the Hackathon Advisor project, selected from local Codex
rollout JSONL files and redacted before publication. The event stream preserves user
requests, assistant messages, tool calls, tool outputs, browser/search events, and
minimal session provenance needed to audit how the project was built.
Privacy filtering
The publisher applied openai/privacy-filter
at revision… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/hackathon-advisor-codex-traces.react-shadcn-codex
React Shadcn Codex Dataset
Description
The React Shadcn Codex is a curated collection of over 3,000 React components that utilize shadcn, Framer Motion, and Lucide React. This dataset provides a valuable resource for developers looking to understand and implement modern React UI components with these popular libraries.
Content
The dataset includes:
3,000+ React components using shadcn UI
Components with Framer Motion animations
Usage examples of Lucide React… See the full description on the dataset page: https://huggingface.co/datasets/valentin-marquez/react-shadcn-codex.gpt-5-codex-250xCodeX-Thinking-Gemma-4-31B-ITAll prompts were taken from Modotte/CodeX-2M-Thinking, which contains multiple traces per prompt whereas this dataset only provides one trace per prompt. Generations were with https://huggingface.co/nvidia/Gemma-4-31B-IT-NVFP4 (a mix of BF16/FP8 weights that NVIDIA configured with FP8 KV cache; benchmarks show performs similarly to BF16 for coding). No system prompt was used.
hemlock-codex-SFT
Hemlock Codex SFT
Supervised fine-tuning dataset for the Hemlock programming language. Contains 552 instruction/output pairs covering algorithms, systems programming, cross-language translation, and practical programs.
Motivation
Benchmark results for hemlang/Hemlock2-Coder-7B (Q8_0, zero-shot) showed weak performance on:
L3 Algorithms (28.6%) — data structures, graph algorithms, DP
L4 Systems Programming (42.9%) — manual memory, concurrency patterns
L5 Translation… See the full description on the dataset page: https://huggingface.co/datasets/hemlang/hemlock-codex-SFT.lemonseed-codex-cogen-heldout
lemonseed-codex-cogen-heldout
LemonSeed — Codex-teacher co-generated heldout set (1h).
Contents
intelligent_codex_heldout_1h.jsonl (28 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
LLM-teacher co-generated instruction/chat data for LemonSeed fine-tuning.
uv-brain-s03_custom_with_rehearsal_v2openthoughts3_numinamath-1.5-pro_mixtureuv-brain-s03_custom_with_rehearsallemonseed-codex-cogen-train
lemonseed-codex-cogen-train
LemonSeed — Codex-teacher co-generated instruction/chat training data (1h).
Contents
intelligent_codex_train_1h.jsonl (84 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
LLM-teacher co-generated instruction/chat data for LemonSeed fine-tuning.
hemlock-codex2-SFT
hemlock-codex2-SFT
v2 of the Hemlock code-generation SFT set: the original 552 codex rows plus an
interpreter-validated expansion targeting the categories the models fail on
(graphs, dp, trees, sorting, search, memory, concurrency, defer,
practical/data-processing), each with codex-style translation variants from all
five source languages (Python, JavaScript, C, Go, Rust).
552 original codex rows
101 new hard generation examples (every one run through the Hemlock
interpreter… See the full description on the dataset page: https://huggingface.co/datasets/hemlang/hemlock-codex2-SFT.balanced_smoltalk_baselemonseed-codex-cogen-expansion
lemonseed-codex-cogen-expansion
LemonSeed — Codex-teacher expansion data, reviewed (v2).
Contents
intelligent_codex_expansion_reviewed_v2.jsonl (336 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
LLM-teacher co-generated instruction/chat data for LemonSeed fine-tuning.
gpt-5-codex-250xLaravel
Laravel Documentation Dataset
Overview
This repository provides a curated dataset focused on the Laravel ecosystem—including Laravel 12.x, Filament 3.x, and several Spatie libraries. The dataset has been designed to support training language models in understanding and generating developer-focused documentation and Q&A content related to these technologies.
Dataset Structure
The dataset is divided into two main files:
laravel_train.jsond:Contains detailed… See the full description on the dataset page: https://huggingface.co/datasets/codeXpedite/Laravel.gpt-5-codex-250xliterary-dataset-pack
Literary Dataset Pack
A rich and diverse multi-task instruction dataset generated from classic public domain literature.
📖 Overview
Literary Dataset Pack is a high-quality instruction-tuning dataset crafted from classic literary texts in the public domain (e.g., Alice in Wonderland). Each paragraph is transformed into multiple supervised tasks designed to train or fine-tune large language models (LLMs) across a wide range of natural language understanding and generation… See the full description on the dataset page: https://huggingface.co/datasets/codeXpedite/literary-dataset-pack.numina_smoltalk_mixture
