Team Ai
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01OpenCoder-LLM /opc-fineweb-code-corpus OpenCoder Dataset The OpenCoder dataset is composed of the following datasets: opc-sft-stage1: the sft data used for opencoder sft-stage1 opc-sft-stage2: the sft data used for opencoder sft-stage2 opc-annealing-corpus: the synthetic data & algorithmic corpus used for opencoder annealing opc-fineweb-code-corpus: the code-related page recalled from fineweb <-- you are here opc-fineweb-math-corpus: the math-related page recalled from finewebrefineCode-code-corpus-meta: the meta-data… See the full description on the dataset page: https://huggingface.co/datasets/OpenCoder-LLM/opc-fineweb-code-corpus.tabular100M<n<1B66 likes7k downloads2y agoHugging Face02OpenCoder-LLM /RefineCode-code-corpus-metaThis dataset consists of meta information (including the repository name and file path) of the raw code data from RefineCode. You can collect those files referring to this metadata and reproduce RefineCode! Note: Currently, we have uploaded the meta data covered by The Stack V2 (About 50% file volume). Due to complex legal considerations, we are unable to provide the complete source code currently. We are working hard to make the remaining part available. RefineCode is a high-quality… See the full description on the dataset page: https://huggingface.co/datasets/OpenCoder-LLM/RefineCode-code-corpus-meta.tabular100M<n<1B32 likes1.3k downloads2y agoHugging Face03SDAIANCAI /Saudilang-Code-Switch-Corpus SCC - Saudilang Code-Switch Corpus The National Center for Artificial Intelligence at the Saudi Data and Artificial Intelligence Authority (SDAIA), published the "SCC" dataset, which stands for "Saudilang Code-Switch Corpus”. This dataset contains a transcription of general conversations taken from a YouTube podcast "Thmanyah" that has been transcribed by the National Center for Artificial Intelligence in SDAIA. The data features three episodes covering different domains: investment… See the full description on the dataset page: https://huggingface.co/datasets/SDAIANCAI/Saudilang-Code-Switch-Corpus.tabularautomatic-speech-recognition1K<n<10K3 likes78 downloads2y agoHugging Face04NickIBrody /powershell-code-corpus PowerShell Code Corpus A dataset of PowerShell scripts collected from public GitHub repositories. Content 4,632 files from popular PowerShell repositories on GitHub Sourced from repos with the highest star counts (quality signal) Each record contains the raw script text plus metadata Fields Field Description source Always github repo owner/repo slug repo_url Full GitHub URL path File path within the repo language Always PowerShell license… See the full description on the dataset page: https://huggingface.co/datasets/NickIBrody/powershell-code-corpus.tabular1K<n<10K0 likes50 downloads6mo agoHugging Face05CatQualia /code-quality-corpusgated CatQualia code-quality corpus — semantic smell classes with before/after fixes 39,383 rows · 25,541,792 bytes · JSON Lines, one object per line. What this is Real code smells paired with the fix: a smell_class that names the semantic problem (not just the syntax), the original lines, the corrected lines, the file and line it came from, and a rationale explaining why the original was wrong. Useful for code-review or repair training where the label has to say what… See the full description on the dataset page: https://huggingface.co/datasets/CatQualia/code-quality-corpus.tabular10K<n<100K0 likes31 downloads26d agoHugging Face06kasuboski /gleam-code-corpus Gleam Code Corpus A structured, attributed corpus of 22,581 Gleam source files collected from 1,528 GitHub repositories for continued pre-training and code generation research. Overview This dataset contains .gleam source files from the Gleam programming language ecosystem. Every file includes full attribution — repo owner, name, URL, license, star count, and Hex.pm package metadata where available. What this is for: Continued pre-training (CPT) of code-focused… See the full description on the dataset page: https://huggingface.co/datasets/kasuboski/gleam-code-corpus.tabulartext-generation10K<n<100K0 likes28 downloads5mo agoHugging Face07NickIBrody /ruby-code-corpus Ruby Code Corpus A large corpus of Ruby source code collected from public GitHub repositories. Content 294,074 files from 3,920+ Ruby repositories on GitHub Sourced from repos ranked by star count (quality signal) Filtered: removed files under 200 bytes (trivial/empty files) Only permissive licenses: MIT, Apache-2.0, BSD, ISC, Ruby Fields Field Description source Always github repo owner/repo slug repo_url Full GitHub URL path File path… See the full description on the dataset page: https://huggingface.co/datasets/NickIBrody/ruby-code-corpus.tabular100K<n<1M0 likes19 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.