Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01OpenCoder-LLM /opc-fineweb-code-corpus OpenCoder Dataset The OpenCoder dataset is composed of the following datasets: opc-sft-stage1: the sft data used for opencoder sft-stage1 opc-sft-stage2: the sft data used for opencoder sft-stage2 opc-annealing-corpus: the synthetic data & algorithmic corpus used for opencoder annealing opc-fineweb-code-corpus: the code-related page recalled from fineweb <-- you are here opc-fineweb-math-corpus: the math-related page recalled from finewebrefineCode-code-corpus-meta: the meta-data… See the full description on the dataset page: https://huggingface.co/datasets/OpenCoder-LLM/opc-fineweb-code-corpus.tabular100M<n<1B65 likes7.2k downloads2y agoHugging Face02rufatronics /clean-code-corpus-v1text1K<n<10K0 likes1.7k downloads2h agoHugging Face03OpenCoder-LLM /RefineCode-code-corpus-metaThis dataset consists of meta information (including the repository name and file path) of the raw code data from RefineCode. You can collect those files referring to this metadata and reproduce RefineCode! Note: Currently, we have uploaded the meta data covered by The Stack V2 (About 50% file volume). Due to complex legal considerations, we are unable to provide the complete source code currently. We are working hard to make the remaining part available. RefineCode is a high-quality… See the full description on the dataset page: https://huggingface.co/datasets/OpenCoder-LLM/RefineCode-code-corpus-meta.tabular100M<n<1B31 likes1.4k downloads2y agoHugging Face04Corpus-NZ /High-Quality-Code High-Quality-Code: Synthetic + Real (MAXIMUM CODE) A massive, high-quality code dataset built with maximum code philosophy – as much code as possible. Components Synthetic syntax-correction dataset – 5M+ examples across 33 languages (original code_syntax_dataset_1GB.csv) Real high-quality code from GitHub – 500 top-starred repositories – BOTH zips and extracted source Current Status: IN PROGRESS Target: 500 repos Currently uploaded: 12 extracted… See the full description on the dataset page: https://huggingface.co/datasets/Corpus-NZ/High-Quality-Code.100M<n<1B0 likes1.3k downloads27d agoHugging Face05teven /code_docstring_corpusHF version of Edinburgh-NLP's Code docstrings corpus text100K<n<1M8 likes319 downloads4y agoHugging Face06nazdef /1gpu-llm-pretraining-corpus-15b-en-it-code 1GPU LLM Pretraining Corpus 15B EN-IT-CODE 1gpu-llm-pretraining-corpus-15b-en-it-code is the canonical document-level pretraining corpus used for the 1GPU LLM family. It was built for training language models from scratch on a mixture of English, Italian and source code. This Hugging Face release contains the clean, deduplicated, document-level corpus. It is intentionally published before tokenization and packing so that the training representation can be deterministically… See the full description on the dataset page: https://huggingface.co/datasets/nazdef/1gpu-llm-pretraining-corpus-15b-en-it-code.text-generation0 likes318 downloads12d agoHugging Face07seongchaeae /sage-code-corpus-v20 likes170 downloads4mo agoHugging Face08archit11 /verl-code-corpus-track-a-file-split archit11/verl-code-corpus-track-a-file-split Repository-specific code corpus extracted from the verl project and split by file for training/evaluation. What is in this dataset Source corpus: data/code_corpus_verl Total files: 214 Train files: 172 Validation files: 21 Test files: 21 File type filter: .py Split mode: file (file-level holdout) Each row has: file_name: flattened source file name text: full file contents Training context This dataset was used… See the full description on the dataset page: https://huggingface.co/datasets/archit11/verl-code-corpus-track-a-file-split.texttext-generationn<1K0 likes104 downloads8mo agoHugging Face09burtenshaw /PleIAs_common_corpus_code_classificationtext100K<n<1M1 likes95 downloads1y agoHugging Face10DatarrX /Burmese-English-Code-Mixed-Corpus 🇲🇲 Burmese-English Code-Mixed Corpus ꒰ 1,111 Rows ꒱ A high-quality, human-curated dataset of code-mixed Burmese and English sentences, specifically designed for Natural Language Processing (NLP) and Machine Learning (ML) research. Dataset Details Organization: DatarrX Creator: Khant Sint Heinn (Kalix Louis) Number of Rows: 1,111 Language: Burmese (Unicode) & English Mix Dataset Format: .txt License: Apache 2.0 Description The Burmese-English Code-Mixed… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/Burmese-English-Code-Mixed-Corpus.texttext-generation1K<n<10K9 likes81 downloads6mo agoHugging Face11Mir-2002 /python_code_docstring_ast_corpus Overview This dataset contains 34,000+ rows of code-docstring-ast data along with additional metadata. Data was gathered from various Python libraries and frameworks and their publicly available GitHub repos. This dataset was created for the purpose of training the CodeT5+ transformer on AST-enhanced code-to-doc tasks. Sources The dataset was gathered from various GitHub repos sampled from this repo by Vinta. The 26 repos are: matplotlib pytorch cryptography django… See the full description on the dataset page: https://huggingface.co/datasets/Mir-2002/python_code_docstring_ast_corpus.textsummarization10K<n<100K1 likes66 downloads1y agoHugging Face12SDAIANCAI /Saudilang-Code-Switch-Corpus SCC - Saudilang Code-Switch Corpus The National Center for Artificial Intelligence at the Saudi Data and Artificial Intelligence Authority (SDAIA), published the "SCC" dataset, which stands for "Saudilang Code-Switch Corpus”. This dataset contains a transcription of general conversations taken from a YouTube podcast "Thmanyah" that has been transcribed by the National Center for Artificial Intelligence in SDAIA. The data features three episodes covering different domains: investment… See the full description on the dataset page: https://huggingface.co/datasets/SDAIANCAI/Saudilang-Code-Switch-Corpus.tabularautomatic-speech-recognition1K<n<10K3 likes59 downloads2y agoHugging Face13krystv /code-corpus-llm-training Code Corpus for LLM Training Manually collected from top open-source repositories across: video/graphics editors, browsers, terminals, UI/UX, Qt/QML, Flutter, Rust, Python, ethical hacking, system-level, game engines, web frameworks, and more. Stats Records: 240,378 Raw text: 2,156,908,643 chars (~2.01 GB) Domains: 20 Domains web_ui: 32,354 records cpp: 29,792 records kotlin_android: 19,476 records ui_ux_design: 19,382 records rust: 15,440 records python:… See the full description on the dataset page: https://huggingface.co/datasets/krystv/code-corpus-llm-training.text100K<n<1M0 likes55 downloads5mo agoHugging Face14ashleymendoza /code-corpus Code Video Text Data Notes Dataset summary Preparation notes and schema examples for Code tasks using Video Text data. Full source material is intentionally not bundled, so provenance and licensing remain explicit. Included material prepare.py — loading, cleaning, and split preparation code. dataset_infos.json — schema and split metadata. metadata_sample.jsonl — small, human-readable records for checking the schema. README.md — data card and usage… See the full description on the dataset page: https://huggingface.co/datasets/ashleymendoza/code-corpus.0 likes53 downloads28d agoHugging Face15nazdef /1gpu-llm-validation-corpus-en-it-code 1GPU LLM Validation Corpus (EN/IT/Code, document-level) Validation-only corpus for the 1GPU LLM pretraining project (nazdef/1gpu-llm Collection). This repository contains the frozen document-level validation master (T050-SUB002, COMPLETE 2026-09-24), not the packed 48k/ctx2500 training artifact. The packed validation set (~10M tokens) is derived downstream from this master and stays local. Relationship to the training corpus Training corpus:… See the full description on the dataset page: https://huggingface.co/datasets/nazdef/1gpu-llm-validation-corpus-en-it-code.0 likes53 downloads12d agoHugging Face16adraganov /arch-resilience-lpi-260903T1715-code-corpus0 likes52 downloads1mo agoHugging Face17HuggingPanda /SDAIANCAI-Saudilang-Code-Switch-Corpusaudio1K<n<10K2 likes51 downloads2y agoHugging Face18NickIBrody /powershell-code-corpus PowerShell Code Corpus A dataset of PowerShell scripts collected from public GitHub repositories. Content 4,632 files from popular PowerShell repositories on GitHub Sourced from repos with the highest star counts (quality signal) Each record contains the raw script text plus metadata Fields Field Description source Always github repo owner/repo slug repo_url Full GitHub URL path File path within the repo language Always PowerShell license… See the full description on the dataset page: https://huggingface.co/datasets/NickIBrody/powershell-code-corpus.tabular1K<n<10K0 likes47 downloads5mo agoHugging Face19kkomyoeminaung /code-cpt-corpus Dataset Card for code-cpt-corpus Dataset Summary ဒီ dataset က code-cpt-corpus အတွက် ဖန်တီးထားတာပါ။ Languages Myanmar (my) / English (en) Dataset Structure Data Instances { "text": "နမူနာ စာသား", "label": "အညွှန်း" } Data Fields text: main content, label: optional. Data Splits Split Files train data/train.jsonl Licensing Information ဒီ dataset က CC BY-NC… See the full description on the dataset page: https://huggingface.co/datasets/kkomyoeminaung/code-cpt-corpus.text100K<n<1M0 likes45 downloads3mo agoHugging Face20VishaalY /synthetic-code-corpustextquestion-answering100K<n<1M0 likes41 downloads2y agoHugging Face21archit11 /hyperswitch-code-corpus-track-a archit11/hyperswitch-code-corpus-track-a Repository-specific code corpus extracted from hyperswitch and split by file for training/evaluation. What is in this dataset Source corpus: data/code_corpus_hyperswitch Total files: 300 Train files: 270 Validation files: 30 Test files: 0 File type filter: .rs Split mode: file (file-level holdout) Each row has: file_name: flattened source file name text: full file contents Training context This dataset was used for… See the full description on the dataset page: https://huggingface.co/datasets/archit11/hyperswitch-code-corpus-track-a.texttext-generationn<1K0 likes36 downloads8mo agoHugging Face22rigidhat /construction-code-corpus-v1 Construction Code-Citation Corpus v1 Open dataset of construction-site incident narratives paired with OIICS hazard codes (event, source, nature, body) and OSHA 29 CFR 1926 citation candidates. Built for the Adaption Labs AutoScientist Challenge ("All Other Domains" category). Sources OSHA Severe Injury Reports (DOL, public domain): 2015-01 → 2025-08, 103,750 records. Each row has Final Narrative (incident text) plus OIICS classification codes for the event… See the full description on the dataset page: https://huggingface.co/datasets/rigidhat/construction-code-corpus-v1.10K<n<100K0 likes34 downloads4mo agoHugging Face23seongchaeae /sage-code-corpus-v1text1M<n<10M1 likes31 downloads4mo agoHugging Face24Collaops /code-agent-corpus Synthetic Code-Agent Run Corpus (Projected) ⚠️ SYNTHETIC DATA — NOT REAL PRODUCTION TELEMETRY. This dataset is synthetically generated for capacity planning / illustration of a code-fixing agent platform's data pipeline. It does not contain real user data, real repositories, real credentials, or real production captures. Identifiers and entity names are anonymized (repo-A…repo-E) and run IDs/timestamps are fabricated. What this is A representative, labeled… See the full description on the dataset page: https://huggingface.co/datasets/Collaops/code-agent-corpus.text-generation1K<n<10K0 likes30 downloads4mo agoHugging Face25coding-assistant-custom /mini-code-corpus Dataset Card for "mini-code-corpus" More Information needed textn<1K1 likes29 downloads3y agoHugging Face26CatQualia /code-quality-corpusgated CatQualia code-quality corpus — semantic smell classes with before/after fixes 39,383 rows · 25,541,792 bytes · JSON Lines, one object per line. What this is Real code smells paired with the fix: a smell_class that names the semantic problem (not just the syntax), the original lines, the corrected lines, the file and line it came from, and a rationale explaining why the original was wrong. Useful for code-review or repair training where the label has to say what… See the full description on the dataset page: https://huggingface.co/datasets/CatQualia/code-quality-corpus.tabular10K<n<100K0 likes28 downloads23d agoHugging Face27kasuboski /gleam-code-corpus Gleam Code Corpus A structured, attributed corpus of 22,581 Gleam source files collected from 1,528 GitHub repositories for continued pre-training and code generation research. Overview This dataset contains .gleam source files from the Gleam programming language ecosystem. Every file includes full attribution — repo owner, name, URL, license, star count, and Hex.pm package metadata where available. What this is for: Continued pre-training (CPT) of code-focused… See the full description on the dataset page: https://huggingface.co/datasets/kasuboski/gleam-code-corpus.tabulartext-generation10K<n<100K0 likes27 downloads4mo agoHugging Face28Corpus-NZ /Code-Syntax-Expanded Code-Syntax-Expanded A massive, high-quality synthetic dataset for training LLMs to identify and correct syntax errors across 33 programming languages. Contains 5+ million unique examples (~1.1 GB) with English explanations – no artificial padding, no duplicate rows. 📊 Dataset Overview Property Value Total rows 5,000,000+ File size ~1.1 GB (uncompressed CSV) Languages 33 Unique templates 160+ error patterns Format CSV (4 columns) License… See the full description on the dataset page: https://huggingface.co/datasets/Corpus-NZ/Code-Syntax-Expanded.text10M<n<100M0 likes24 downloads1mo agoHugging Face29archit11 /verl-code-corpustextn<1K0 likes23 downloads8mo agoHugging Face30ostapeno /opc-annealing-corpus-synth-qa-code_python_js_tstext1M<n<10M0 likes20 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.