Team Ai
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01archit11 /verl-code-corpus-track-a-file-split archit11/verl-code-corpus-track-a-file-split Repository-specific code corpus extracted from the verl project and split by file for training/evaluation. What is in this dataset Source corpus: data/code_corpus_verl Total files: 214 Train files: 172 Validation files: 21 Test files: 21 File type filter: .py Split mode: file (file-level holdout) Each row has: file_name: flattened source file name text: full file contents Training context This dataset was used… See the full description on the dataset page: https://huggingface.co/datasets/archit11/verl-code-corpus-track-a-file-split.texttext-generationn<1K0 likes108 downloads8mo agoHugging Face02DatarrX /Burmese-English-Code-Mixed-Corpus 🇲🇲 Burmese-English Code-Mixed Corpus ꒰ 1,111 Rows ꒱ A high-quality, human-curated dataset of code-mixed Burmese and English sentences, specifically designed for Natural Language Processing (NLP) and Machine Learning (ML) research. Dataset Details Organization: DatarrX Creator: Khant Sint Heinn (Kalix Louis) Number of Rows: 1,111 Language: Burmese (Unicode) & English Mix Dataset Format: .txt License: Apache 2.0 Description The Burmese-English Code-Mixed… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/Burmese-English-Code-Mixed-Corpus.texttext-generation1K<n<10K9 likes82 downloads6mo agoHugging Face03Mir-2002 /python_code_docstring_ast_corpus Overview This dataset contains 34,000+ rows of code-docstring-ast data along with additional metadata. Data was gathered from various Python libraries and frameworks and their publicly available GitHub repos. This dataset was created for the purpose of training the CodeT5+ transformer on AST-enhanced code-to-doc tasks. Sources The dataset was gathered from various GitHub repos sampled from this repo by Vinta. The 26 repos are: matplotlib pytorch cryptography django… See the full description on the dataset page: https://huggingface.co/datasets/Mir-2002/python_code_docstring_ast_corpus.textsummarization10K<n<100K1 likes52 downloads1y agoHugging Face04archit11 /hyperswitch-code-corpus-track-a archit11/hyperswitch-code-corpus-track-a Repository-specific code corpus extracted from hyperswitch and split by file for training/evaluation. What is in this dataset Source corpus: data/code_corpus_hyperswitch Total files: 300 Train files: 270 Validation files: 30 Test files: 0 File type filter: .rs Split mode: file (file-level holdout) Each row has: file_name: flattened source file name text: full file contents Training context This dataset was used for… See the full description on the dataset page: https://huggingface.co/datasets/archit11/hyperswitch-code-corpus-track-a.texttext-generationn<1K0 likes41 downloads8mo agoHugging Face05Collaops /code-agent-corpus Synthetic Code-Agent Run Corpus (Projected) ⚠️ SYNTHETIC DATA — NOT REAL PRODUCTION TELEMETRY. This dataset is synthetically generated for capacity planning / illustration of a code-fixing agent platform's data pipeline. It does not contain real user data, real repositories, real credentials, or real production captures. Identifiers and entity names are anonymized (repo-A…repo-E) and run IDs/timestamps are fabricated. What this is A representative, labeled… See the full description on the dataset page: https://huggingface.co/datasets/Collaops/code-agent-corpus.text-generation1K<n<10K0 likes31 downloads4mo agoHugging Face06kasuboski /gleam-code-corpus Gleam Code Corpus A structured, attributed corpus of 22,581 Gleam source files collected from 1,528 GitHub repositories for continued pre-training and code generation research. Overview This dataset contains .gleam source files from the Gleam programming language ecosystem. Every file includes full attribution — repo owner, name, URL, license, star count, and Hex.pm package metadata where available. What this is for: Continued pre-training (CPT) of code-focused… See the full description on the dataset page: https://huggingface.co/datasets/kasuboski/gleam-code-corpus.tabulartext-generation10K<n<100K0 likes28 downloads5mo agoHugging Face07hksamm /Burmese-English-Code-Mixed-Corpus 🇲🇲 Burmese-English Code-Mixed Corpus ꒰ 1,111 Rows ꒱ A high-quality, human-curated dataset of code-mixed Burmese and English sentences, specifically designed for Natural Language Processing (NLP) and Machine Learning (ML) research. Dataset Details Organization: DatarrX Creator: Khant Sint Heinn (Kalix Louis) Number of Rows: 1,111 Language: Burmese (Unicode) & English Mix Dataset Format: .txt License: Apache 2.0 Description The Burmese-English Code-Mixed… See the full description on the dataset page: https://huggingface.co/datasets/hksamm/Burmese-English-Code-Mixed-Corpus.texttext-generation1K<n<10K0 likes17 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.