datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GPT-5.5-CodexThis dataset was generated using teich by TeichAI
GPT-5.5 Agent traces
This directory contains raw agent trace files generated by teich.
JSONL files: 317
Model metadata: gpt-5.5
Training-ready tools
Generated agent traces carry configured or recovered tool schemas so tools remain available for training even when a session did not call them.
Native Claude Code imports recover schemas for Claude Code and Claude Desktop built-ins, plus conservative name-derived MCP… See the full description on the dataset page: https://huggingface.co/datasets/AletheiaResearch/GPT-5.5-Codex.code_x_glue_cc_code_completion_token
Dataset Card for "code_x_glue_cc_code_completion_token"
Dataset Summary
CodeXGLUE CodeCompletion-token dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Code/CodeCompletion-token
Predict next code token given context of previous tokens. Models are evaluated by token level accuracy.
Code completion is a one of the most widely used features in software development through IDEs. An effective code completion tool could improve software… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_cc_code_completion_token.CodeX-2M-Thinking
Modotte
Note: This dataset is part of the lineup CodeX by Modotte. You can get lots of datasets in this same lineup, with the main focus on providing very high-quality datasets for model training and fine-tuning.
This dataset is fully synthetic, curated from high-quality public sources and enhanced with synthetic data generated using both closed and open-source models. It serves as a strong foundation for instruction-based model tuning and fine-tuning, offering one of the… See the full description on the dataset page: https://huggingface.co/datasets/Modotte/CodeX-2M-Thinking.code_x_glue_cc_code_completion_line
Dataset Card for "code_x_glue_cc_code_completion_line"
Dataset Summary
CodeXGLUE CodeCompletion-line dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Code/CodeCompletion-line
Complete the unfinished line given previous context. Models are evaluated by exact match and edit similarity.
We propose line completion task to test model's ability to autocomplete a line. Majority code completion systems behave well in token level completion, but fail in… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_cc_code_completion_line.aleppo-codex
NuBerea Aleppo Codex
Hebrew text of the Aleppo Codex, the oldest near-complete manuscript of the Masoretic Text (circa 10th century) and widely considered the most authoritative Hebrew Bible manuscript. Old Testament only. This dataset is part of the NuBerea curated corpus estate of biblical and patristic source texts.
License and Attribution
This dataset is licensed CC-BY-4.0. The Aleppo Codex text itself is public domain; the digitization used here is released… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/aleppo-codex.CodeX-2M-Thinking
Modotte
Note: This dataset is part of the lineup CodeX by Modotte. You can get lots of datasets in this same lineup, with the main focus on providing very high-quality datasets for model training and fine-tuning.
This dataset is fully synthetic, curated from high-quality public sources and enhanced with synthetic data generated using both closed and open-source models. It serves as a strong foundation for instruction-based model tuning and fine-tuning, offering one of the… See the full description on the dataset page: https://huggingface.co/datasets/txchmechanicus/CodeX-2M-Thinking.CodeX-7M-Non-Thinking
Modotte
Note: This dataset is part of the lineup CodeX by Modotte. You can get lots of datasets in this same lineup, with the main focus on providing very high-quality datasets for model training and fine-tuning.
This dataset is curated from high-quality public sources and enhanced with synthetic data from both closed and open-source models. It serves as a strong foundation for instruction-based model tuning and fine-tuning, offering one of the most refined and extensive… See the full description on the dataset page: https://huggingface.co/datasets/Modotte/CodeX-7M-Non-Thinking.code_x_glue_cc_cloze_testing_all
Dataset Card for "code_x_glue_cc_cloze_testing_all"
Dataset Summary
CodeXGLUE ClozeTesting-all dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Code/ClozeTesting-all
Cloze tests are widely adopted in Natural Languages Processing to evaluate the performance of the trained language models. The task is aimed to predict the answers for the blank with the context of the blank, which can be formulated as a multi-choice classification problem.
Here we… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_cc_cloze_testing_all.CodeX-2M-Thinking
Modotte
Note: This dataset is part of the lineup CodeX by Modotte. You can get lots of datasets in this same lineup, with the main focus on providing very high-quality datasets for model training and fine-tuning.
This dataset is fully synthetic, curated from high-quality public sources and enhanced with synthetic data generated using both closed and open-source models. It serves as a strong foundation for instruction-based model tuning and fine-tuning, offering one of… See the full description on the dataset page: https://huggingface.co/datasets/me-aas/CodeX-2M-Thinking.code_x_glue_cc_cloze_testing_maxmin
Dataset Card for "code_x_glue_cc_cloze_testing_maxmin"
Dataset Summary
CodeXGLUE ClozeTesting-maxmin dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Code/ClozeTesting-maxmin
Cloze tests are widely adopted in Natural Languages Processing to evaluate the performance of the trained language models. The task is aimed to predict the answers for the blank with the context of the blank, which can be formulated as a multi-choice classification problem.… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_cc_cloze_testing_maxmin.share-codex
Local Coding Agent Sessions
Local coding-agent session transcripts exported with
share-codex.
Rows are in train.jsonl. Each row contains id, ordered messages, and
metadata. Messages include user prompts, assistant responses/tool calls, and
tool outputs. Internal prompts and reasoning are excluded by default.
Dataset Statistics
Statistics are computed from the final merged train.jsonl rows at export time.
Metric
Value
Sessions
4019
User turns
16639… See the full description on the dataset page: https://huggingface.co/datasets/nmuendler/share-codex.Kimi-K3-CodexThis dataset was generated using teich by TeichAI
Kimi-K3 Codex traces
This directory contains raw agent trace files generated by teich.
JSONL files: 5
Model metadata: moonshotai/kimi-k3
Training-ready tools
Generated agent traces carry configured or recovered tool schemas so tools remain available for training even when a session did not call them.
Native Claude Code imports recover schemas for Claude Code and Claude Desktop built-ins, plus conservative… See the full description on the dataset page: https://huggingface.co/datasets/AletheiaResearch/Kimi-K3-Codex.CodeX-2M-Thinking
Modotte
Note: This dataset is part of the lineup CodeX by Modotte. You can get lots of datasets in this same lineup, with the main focus on providing very high-quality datasets for model training and fine-tuning.
This dataset is fully synthetic, curated from high-quality public sources and enhanced with synthetic data generated using both closed and open-source models. It serves as a strong foundation for instruction-based model tuning and fine-tuning, offering one of the… See the full description on the dataset page: https://huggingface.co/datasets/adrianmele/CodeX-2M-Thinking.my-personal-codex-data
Coding Agent Conversation Logs
This is a performance art project. Anthropic built their models on the world's freely shared information, then introduced increasingly dystopian data policies to stop anyone else from doing the same with their data — pulling up the ladder behind them. DataClaw lets you throw the ladder back down. The dataset it produces is yours to share.
Exported with DataClaw.
Tag: dataclaw — Browse all DataClaw datasets
Stats
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/xuechengjiang/my-personal-codex-data.codex-math-qaSolution by codex-davinci-002 for math_qamy-personal-codex-data
Coding Agent Conversation Logs
This is a performance art project. Anthropic built their models on the world's freely shared information, then introduced increasingly dystopian data policies to stop anyone else from doing the same with their data — pulling up the ladder behind them. DataClaw lets you throw the ladder back down. The dataset it produces is yours to share.
Exported with DataClaw.
Tag: dataclaw — Browse all DataClaw datasets
Stats
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/michaelwaves/my-personal-codex-data.my-personal-codex-data
Coding Agent Conversation Logs
This is a performance art project. Anthropic built their models on the world's freely shared information, then introduced increasingly dystopian data policies to stop anyone else from doing the same with their data - pulling up the ladder behind them. DataClaw lets you throw the ladder back down. The dataset it produces is yours to share.
Exported with DataClaw.
Tag: dataclaw - Browse all DataClaw datasets
Stats
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/wop/my-personal-codex-data.my-personal-codex-data
Coding Agent Conversation Logs
This is a performance art project. Anthropic built their models on the world's freely shared information, then introduced increasingly dystopian data policies to stop anyone else from doing the same with their data — pulling up the ladder behind them. DataClaw lets you throw the ladder back down. The dataset it produces is yours to share.
Exported with DataClaw.
Tag: dataclaw — Browse all DataClaw datasets
Stats
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/akenove/my-personal-codex-data.codex-opencomic-dev
OpenComic-Continue Corpus
This versioned research dataset is intended for leakage-safe comic understanding and continuation. Each example retains work/series/page identity and a machine-readable provenance link.
Sources and licences
Only per-work records accepted by data/licenses/license_manifest.parquet may enter the core split. Public Domain, CC0, and CC BY are the default classes. Source-specific licence and attribution metadata govern every item; this card… See the full description on the dataset page: https://huggingface.co/datasets/sagnik-mukherjee/codex-opencomic-dev.my-personal-codex-data
Coding Agent Conversation Logs
This is a performance art project. Anthropic built their models on the world's freely shared information, then introduced increasingly dystopian data policies to stop anyone else from doing the same with their data - pulling up the ladder behind them. DataClaw lets you throw the ladder back down. The dataset it produces is yours to share.
Exported with DataClaw.
Tag: dataclaw - Browse all DataClaw datasets
Stats
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/masterda/my-personal-codex-data.my-personal-codex-data
Coding Agent Conversation Logs
This is a performance art project. Anthropic built their models on the world's freely shared information, then introduced increasingly dystopian data policies to stop anyone else from doing the same with their data - pulling up the ladder behind them. DataClaw lets you throw the ladder back down. The dataset it produces is yours to share.
Exported with DataClaw.
Tag: dataclaw - Browse all DataClaw datasets
Stats
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/sinhaankur/my-personal-codex-data.my-personal-codex-data
Coding Agent Conversation Logs
This is a performance art project. Anthropic built their models on the world's freely shared information, then introduced increasingly dystopian data policies to stop anyone else from doing the same with their data - pulling up the ladder behind them. DataClaw lets you throw the ladder back down. The dataset it produces is yours to share.
Exported with DataClaw.
Tag: dataclaw - Browse all DataClaw datasets
Stats
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/wuuski/my-personal-codex-data.hackathon-advisor-codex-traces
Hackathon Advisor Codex Session Traces
Real Codex session logs for the Hackathon Advisor project, selected from local Codex
rollout JSONL files and redacted before publication. The event stream preserves user
requests, assistant messages, tool calls, tool outputs, browser/search events, and
minimal session provenance needed to audit how the project was built.
Privacy filtering
The publisher applied openai/privacy-filter
at revision… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/hackathon-advisor-codex-traces.react-shadcn-codex
React Shadcn Codex Dataset
Description
The React Shadcn Codex is a curated collection of over 3,000 React components that utilize shadcn, Framer Motion, and Lucide React. This dataset provides a valuable resource for developers looking to understand and implement modern React UI components with these popular libraries.
Content
The dataset includes:
3,000+ React components using shadcn UI
Components with Framer Motion animations
Usage examples of Lucide React… See the full description on the dataset page: https://huggingface.co/datasets/valentin-marquez/react-shadcn-codex.gpt-5-codex-250xcodex5.5_ToT
Codex 5.5 Handwritten TropicalGT/ToricGT ToT Reasoning
This dataset contains individually authored Tree-of-Thought style reasoning records for TropicalGT/ToricGT training. Each row represents one accepted JSON reasoning artifact, flattened into stable Parquet columns for filtering and dataset-viewer compatibility while preserving the complete canonical JSON object in record_json.
The source dataset is maintained in tot_reasoning_dataset_handwritten/. The rejected bulk-generated… See the full description on the dataset page: https://huggingface.co/datasets/AmelieSchreiber/codex5.5_ToT.method2test_v2024hemlock-codex-SFT
Hemlock Codex SFT
Supervised fine-tuning dataset for the Hemlock programming language. Contains 552 instruction/output pairs covering algorithms, systems programming, cross-language translation, and practical programs.
Motivation
Benchmark results for hemlang/Hemlock2-Coder-7B (Q8_0, zero-shot) showed weak performance on:
L3 Algorithms (28.6%) — data structures, graph algorithms, DP
L4 Systems Programming (42.9%) — manual memory, concurrency patterns
L5 Translation… See the full description on the dataset page: https://huggingface.co/datasets/hemlang/hemlock-codex-SFT.Codex-Reasoning-4000x-filteredNote: This dataset is part of the lineup with reasoning. You can get lots of datasets in this same lineup, with the main focus on providing very high-quality datasets for model training and fine-tuning.
This dataset is fully synthetic, curated from high-quality public sources and enhanced with synthetic data generated using both closed and open-source models. It serves as a strong foundation for instruction-based model tuning and fine-tuning, offering one of the most refined and extensive… See the full description on the dataset page: https://huggingface.co/datasets/nphearum/Codex-Reasoning-4000x-filtered.hemlock-codex3-SFT
hemlock-codex3-SFT
Code-generation SFT data for the Hemlock
programming language: translation tasks (C/Go/JavaScript/Python/Rust → Hemlock),
algorithm/systems generation, and stdlib recall — every reference answer
execution-verified against the Hemlock interpreter at build time.
Supersedes and merges hemlock-codex-SFT
(a strict subset of codex2), hemlock-codex2-SFT,
and hemlock-formulary-SFT.
What changed vs codex2/formulary
Fenced outputs. Every output is a… See the full description on the dataset page: https://huggingface.co/datasets/hemlang/hemlock-codex3-SFT.
