Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AletheiaResearch /GPT-5.5-CodexThis dataset was generated using teich by TeichAI GPT-5.5 Agent traces This directory contains raw agent trace files generated by teich. JSONL files: 317 Model metadata: gpt-5.5 Training-ready tools Generated agent traces carry configured or recovered tool schemas so tools remain available for training even when a session did not call them. Native Claude Code imports recover schemas for Claude Code and Claude Desktop built-ins, plus conservative name-derived MCP… See the full description on the dataset page: https://huggingface.co/datasets/AletheiaResearch/GPT-5.5-Codex.tabulartext-generationn<1K14 likes1.6k downloads3mo agoHugging Face02google /code_x_glue_cc_code_completion_token Dataset Card for "code_x_glue_cc_code_completion_token" Dataset Summary CodeXGLUE CodeCompletion-token dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Code/CodeCompletion-token Predict next code token given context of previous tokens. Models are evaluated by token level accuracy. Code completion is a one of the most widely used features in software development through IDEs. An effective code completion tool could improve software… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_cc_code_completion_token.texttext-generation100K<n<1M15 likes1.1k downloads3y agoHugging Face03Modotte /CodeX-2M-Thinking Modotte Note: This dataset is part of the lineup CodeX by Modotte. You can get lots of datasets in this same lineup, with the main focus on providing very high-quality datasets for model training and fine-tuning. This dataset is fully synthetic, curated from high-quality public sources and enhanced with synthetic data generated using both closed and open-source models. It serves as a strong foundation for instruction-based model tuning and fine-tuning, offering one of the… See the full description on the dataset page: https://huggingface.co/datasets/Modotte/CodeX-2M-Thinking.texttext-generation1M<n<10M129 likes834 downloads8mo agoHugging Face04google /code_x_glue_cc_code_completion_line Dataset Card for "code_x_glue_cc_code_completion_line" Dataset Summary CodeXGLUE CodeCompletion-line dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Code/CodeCompletion-line Complete the unfinished line given previous context. Models are evaluated by exact match and edit similarity. We propose line completion task to test model's ability to autocomplete a line. Majority code completion systems behave well in token level completion, but fail in… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_cc_code_completion_line.texttext-generation10K<n<100K10 likes457 downloads3y agoHugging Face05NuBerea /aleppo-codexgated NuBerea Aleppo Codex Hebrew text of the Aleppo Codex, the oldest near-complete manuscript of the Masoretic Text (circa 10th century) and widely considered the most authoritative Hebrew Bible manuscript. Old Testament only. This dataset is part of the NuBerea curated corpus estate of biblical and patristic source texts. License and Attribution This dataset is licensed CC-BY-4.0. The Aleppo Codex text itself is public domain; the digitization used here is released… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/aleppo-codex.tabulartext-generation10K<n<100K0 likes446 downloads3mo agoHugging Face06txchmechanicus /CodeX-2M-Thinking Modotte Note: This dataset is part of the lineup CodeX by Modotte. You can get lots of datasets in this same lineup, with the main focus on providing very high-quality datasets for model training and fine-tuning. This dataset is fully synthetic, curated from high-quality public sources and enhanced with synthetic data generated using both closed and open-source models. It serves as a strong foundation for instruction-based model tuning and fine-tuning, offering one of the… See the full description on the dataset page: https://huggingface.co/datasets/txchmechanicus/CodeX-2M-Thinking.texttext-generation1M<n<10M0 likes446 downloads5mo agoHugging Face07Modotte /CodeX-7M-Non-Thinking Modotte Note: This dataset is part of the lineup CodeX by Modotte. You can get lots of datasets in this same lineup, with the main focus on providing very high-quality datasets for model training and fine-tuning. This dataset is curated from high-quality public sources and enhanced with synthetic data from both closed and open-source models. It serves as a strong foundation for instruction-based model tuning and fine-tuning, offering one of the most refined and extensive… See the full description on the dataset page: https://huggingface.co/datasets/Modotte/CodeX-7M-Non-Thinking.texttext-generation1M<n<10M25 likes356 downloads8mo agoHugging Face08google /code_x_glue_cc_cloze_testing_all Dataset Card for "code_x_glue_cc_cloze_testing_all" Dataset Summary CodeXGLUE ClozeTesting-all dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Code/ClozeTesting-all Cloze tests are widely adopted in Natural Languages Processing to evaluate the performance of the trained language models. The task is aimed to predict the answers for the blank with the context of the blank, which can be formulated as a multi-choice classification problem. Here we… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_cc_cloze_testing_all.texttext-generation100K<n<1M6 likes290 downloads3y agoHugging Face09me-aas /CodeX-2M-Thinking Modotte Note: This dataset is part of the lineup CodeX by Modotte. You can get lots of datasets in this same lineup, with the main focus on providing very high-quality datasets for model training and fine-tuning. This dataset is fully synthetic, curated from high-quality public sources and enhanced with synthetic data generated using both closed and open-source models. It serves as a strong foundation for instruction-based model tuning and fine-tuning, offering one of… See the full description on the dataset page: https://huggingface.co/datasets/me-aas/CodeX-2M-Thinking.texttext-generation1M<n<10M1 likes281 downloads4mo agoHugging Face10google /code_x_glue_cc_cloze_testing_maxmin Dataset Card for "code_x_glue_cc_cloze_testing_maxmin" Dataset Summary CodeXGLUE ClozeTesting-maxmin dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Code/ClozeTesting-maxmin Cloze tests are widely adopted in Natural Languages Processing to evaluate the performance of the trained language models. The task is aimed to predict the answers for the blank with the context of the blank, which can be formulated as a multi-choice classification problem.… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_cc_cloze_testing_maxmin.texttext-generation1K<n<10K3 likes238 downloads3y agoHugging Face11nmuendler /share-codex Local Coding Agent Sessions Local coding-agent session transcripts exported with share-codex. Rows are in train.jsonl. Each row contains id, ordered messages, and metadata. Messages include user prompts, assistant responses/tool calls, and tool outputs. Internal prompts and reasoning are excluded by default. Dataset Statistics Statistics are computed from the final merged train.jsonl rows at export time. Metric Value Sessions 4019 User turns 16639… See the full description on the dataset page: https://huggingface.co/datasets/nmuendler/share-codex.texttext-generation1K<n<10K8 likes238 downloads1mo agoHugging Face12AletheiaResearch /Kimi-K3-CodexThis dataset was generated using teich by TeichAI Kimi-K3 Codex traces This directory contains raw agent trace files generated by teich. JSONL files: 5 Model metadata: moonshotai/kimi-k3 Training-ready tools Generated agent traces carry configured or recovered tool schemas so tools remain available for training even when a session did not call them. Native Claude Code imports recover schemas for Claude Code and Claude Desktop built-ins, plus conservative… See the full description on the dataset page: https://huggingface.co/datasets/AletheiaResearch/Kimi-K3-Codex.tabulartext-generationn<1K14 likes176 downloads3mo agoHugging Face13adrianmele /CodeX-2M-Thinking Modotte Note: This dataset is part of the lineup CodeX by Modotte. You can get lots of datasets in this same lineup, with the main focus on providing very high-quality datasets for model training and fine-tuning. This dataset is fully synthetic, curated from high-quality public sources and enhanced with synthetic data generated using both closed and open-source models. It serves as a strong foundation for instruction-based model tuning and fine-tuning, offering one of the… See the full description on the dataset page: https://huggingface.co/datasets/adrianmele/CodeX-2M-Thinking.texttext-generation1M<n<10M0 likes146 downloads6mo agoHugging Face14xuechengjiang /my-personal-codex-data Coding Agent Conversation Logs This is a performance art project. Anthropic built their models on the world's freely shared information, then introduced increasingly dystopian data policies to stop anyone else from doing the same with their data — pulling up the ladder behind them. DataClaw lets you throw the ladder back down. The dataset it produces is yours to share. Exported with DataClaw. Tag: dataclaw — Browse all DataClaw datasets Stats Metric Value… See the full description on the dataset page: https://huggingface.co/datasets/xuechengjiang/my-personal-codex-data.texttext-generationn<1K2 likes109 downloads8mo agoHugging Face15theblackcat102 /codex-math-qaSolution by codex-davinci-002 for math_qatexttext-generation10K<n<100K31 likes107 downloads4y agoHugging Face16michaelwaves /my-personal-codex-data Coding Agent Conversation Logs This is a performance art project. Anthropic built their models on the world's freely shared information, then introduced increasingly dystopian data policies to stop anyone else from doing the same with their data — pulling up the ladder behind them. DataClaw lets you throw the ladder back down. The dataset it produces is yours to share. Exported with DataClaw. Tag: dataclaw — Browse all DataClaw datasets Stats Metric Value… See the full description on the dataset page: https://huggingface.co/datasets/michaelwaves/my-personal-codex-data.texttext-generationn<1K1 likes104 downloads8mo agoHugging Face17wop /my-personal-codex-data Coding Agent Conversation Logs This is a performance art project. Anthropic built their models on the world's freely shared information, then introduced increasingly dystopian data policies to stop anyone else from doing the same with their data - pulling up the ladder behind them. DataClaw lets you throw the ladder back down. The dataset it produces is yours to share. Exported with DataClaw. Tag: dataclaw - Browse all DataClaw datasets Stats Metric Value… See the full description on the dataset page: https://huggingface.co/datasets/wop/my-personal-codex-data.texttext-generationn<1K0 likes96 downloads4mo agoHugging Face18akenove /my-personal-codex-data Coding Agent Conversation Logs This is a performance art project. Anthropic built their models on the world's freely shared information, then introduced increasingly dystopian data policies to stop anyone else from doing the same with their data — pulling up the ladder behind them. DataClaw lets you throw the ladder back down. The dataset it produces is yours to share. Exported with DataClaw. Tag: dataclaw — Browse all DataClaw datasets Stats Metric Value… See the full description on the dataset page: https://huggingface.co/datasets/akenove/my-personal-codex-data.texttext-generationn<1K1 likes95 downloads8mo agoHugging Face19sagnik-mukherjee /codex-opencomic-dev OpenComic-Continue Corpus This versioned research dataset is intended for leakage-safe comic understanding and continuation. Each example retains work/series/page identity and a machine-readable provenance link. Sources and licences Only per-work records accepted by data/licenses/license_manifest.parquet may enter the core split. Public Domain, CC0, and CC BY are the default classes. Source-specific licence and attribution metadata govern every item; this card… See the full description on the dataset page: https://huggingface.co/datasets/sagnik-mukherjee/codex-opencomic-dev.imageimage-to-textn<1K0 likes95 downloads2mo agoHugging Face20masterda /my-personal-codex-data Coding Agent Conversation Logs This is a performance art project. Anthropic built their models on the world's freely shared information, then introduced increasingly dystopian data policies to stop anyone else from doing the same with their data - pulling up the ladder behind them. DataClaw lets you throw the ladder back down. The dataset it produces is yours to share. Exported with DataClaw. Tag: dataclaw - Browse all DataClaw datasets Stats Metric Value… See the full description on the dataset page: https://huggingface.co/datasets/masterda/my-personal-codex-data.texttext-generationn<1K0 likes91 downloads1mo agoHugging Face21sinhaankur /my-personal-codex-data Coding Agent Conversation Logs This is a performance art project. Anthropic built their models on the world's freely shared information, then introduced increasingly dystopian data policies to stop anyone else from doing the same with their data - pulling up the ladder behind them. DataClaw lets you throw the ladder back down. The dataset it produces is yours to share. Exported with DataClaw. Tag: dataclaw - Browse all DataClaw datasets Stats Metric Value… See the full description on the dataset page: https://huggingface.co/datasets/sinhaankur/my-personal-codex-data.texttext-generationn<1K0 likes82 downloads5mo agoHugging Face22wuuski /my-personal-codex-data Coding Agent Conversation Logs This is a performance art project. Anthropic built their models on the world's freely shared information, then introduced increasingly dystopian data policies to stop anyone else from doing the same with their data - pulling up the ladder behind them. DataClaw lets you throw the ladder back down. The dataset it produces is yours to share. Exported with DataClaw. Tag: dataclaw - Browse all DataClaw datasets Stats Metric Value… See the full description on the dataset page: https://huggingface.co/datasets/wuuski/my-personal-codex-data.texttext-generationn<1K0 likes77 downloads6mo agoHugging Face23build-small-hackathon /hackathon-advisor-codex-traces Hackathon Advisor Codex Session Traces Real Codex session logs for the Hackathon Advisor project, selected from local Codex rollout JSONL files and redacted before publication. The event stream preserves user requests, assistant messages, tool calls, tool outputs, browser/search events, and minimal session provenance needed to audit how the project was built. Privacy filtering The publisher applied openai/privacy-filter at revision… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/hackathon-advisor-codex-traces.tabulartext-generationn<1K0 likes77 downloads4mo agoHugging Face24valentin-marquez /react-shadcn-codex React Shadcn Codex Dataset Description The React Shadcn Codex is a curated collection of over 3,000 React components that utilize shadcn, Framer Motion, and Lucide React. This dataset provides a valuable resource for developers looking to understand and implement modern React UI components with these popular libraries. Content The dataset includes: 3,000+ React components using shadcn UI Components with Framer Motion animations Usage examples of Lucide React… See the full description on the dataset page: https://huggingface.co/datasets/valentin-marquez/react-shadcn-codex.texttext-generation1K<n<10K13 likes67 downloads2y agoHugging Face25TeichAI /gpt-5-codex-250xtexttext-generationn<1K14 likes66 downloads11mo agoHugging Face26AmelieSchreiber /codex5.5_ToT Codex 5.5 Handwritten TropicalGT/ToricGT ToT Reasoning This dataset contains individually authored Tree-of-Thought style reasoning records for TropicalGT/ToricGT training. Each row represents one accepted JSON reasoning artifact, flattened into stable Parquet columns for filtering and dataset-viewer compatibility while preserving the complete canonical JSON object in record_json. The source dataset is maintained in tot_reasoning_dataset_handwritten/. The rejected bulk-generated… See the full description on the dataset page: https://huggingface.co/datasets/AmelieSchreiber/codex5.5_ToT.tabulartext-generation1K<n<10K1 likes65 downloads4mo agoHugging Face27CodexAI /method2test_v2024texttext-generation100K<n<1M0 likes57 downloads2y agoHugging Face28hemlang /hemlock-codex-SFT Hemlock Codex SFT Supervised fine-tuning dataset for the Hemlock programming language. Contains 552 instruction/output pairs covering algorithms, systems programming, cross-language translation, and practical programs. Motivation Benchmark results for hemlang/Hemlock2-Coder-7B (Q8_0, zero-shot) showed weak performance on: L3 Algorithms (28.6%) — data structures, graph algorithms, DP L4 Systems Programming (42.9%) — manual memory, concurrency patterns L5 Translation… See the full description on the dataset page: https://huggingface.co/datasets/hemlang/hemlock-codex-SFT.texttext-generationn<1K1 likes51 downloads6mo agoHugging Face29nphearum /Codex-Reasoning-4000x-filteredNote: This dataset is part of the lineup with reasoning. You can get lots of datasets in this same lineup, with the main focus on providing very high-quality datasets for model training and fine-tuning. This dataset is fully synthetic, curated from high-quality public sources and enhanced with synthetic data generated using both closed and open-source models. It serves as a strong foundation for instruction-based model tuning and fine-tuning, offering one of the most refined and extensive… See the full description on the dataset page: https://huggingface.co/datasets/nphearum/Codex-Reasoning-4000x-filtered.texttext-generation10K<n<100K1 likes44 downloads6mo agoHugging Face30hemlang /hemlock-codex3-SFT hemlock-codex3-SFT Code-generation SFT data for the Hemlock programming language: translation tasks (C/Go/JavaScript/Python/Rust → Hemlock), algorithm/systems generation, and stdlib recall — every reference answer execution-verified against the Hemlock interpreter at build time. Supersedes and merges hemlock-codex-SFT (a strict subset of codex2), hemlock-codex2-SFT, and hemlock-formulary-SFT. What changed vs codex2/formulary Fenced outputs. Every output is a… See the full description on the dataset page: https://huggingface.co/datasets/hemlang/hemlock-codex3-SFT.texttext-generation1K<n<10K0 likes35 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.