datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
verl-code-corpus-track-a-file-split
archit11/verl-code-corpus-track-a-file-split
Repository-specific code corpus extracted from the verl project and split by file for training/evaluation.
What is in this dataset
Source corpus: data/code_corpus_verl
Total files: 214
Train files: 172
Validation files: 21
Test files: 21
File type filter: .py
Split mode: file (file-level holdout)
Each row has:
file_name: flattened source file name
text: full file contents
Training context
This dataset was used… See the full description on the dataset page: https://huggingface.co/datasets/archit11/verl-code-corpus-track-a-file-split.Burmese-English-Code-Mixed-Corpus
🇲🇲 Burmese-English Code-Mixed Corpus ꒰ 1,111 Rows ꒱
A high-quality, human-curated dataset of code-mixed Burmese and English sentences, specifically designed for Natural Language Processing (NLP) and Machine Learning (ML) research.
Dataset Details
Organization: DatarrX
Creator: Khant Sint Heinn (Kalix Louis)
Number of Rows: 1,111
Language: Burmese (Unicode) & English Mix
Dataset Format: .txt
License: Apache 2.0
Description
The Burmese-English Code-Mixed… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/Burmese-English-Code-Mixed-Corpus.python_code_docstring_ast_corpus
Overview
This dataset contains 34,000+ rows of code-docstring-ast data along with additional metadata. Data was gathered from various Python libraries and frameworks and their
publicly available GitHub repos. This dataset was created for the purpose of training the CodeT5+ transformer on AST-enhanced code-to-doc tasks.
Sources
The dataset was gathered from various GitHub repos sampled from this repo by Vinta.
The 26 repos are:
matplotlib
pytorch
cryptography
django… See the full description on the dataset page: https://huggingface.co/datasets/Mir-2002/python_code_docstring_ast_corpus.hyperswitch-code-corpus-track-a
archit11/hyperswitch-code-corpus-track-a
Repository-specific code corpus extracted from hyperswitch and split by file for training/evaluation.
What is in this dataset
Source corpus: data/code_corpus_hyperswitch
Total files: 300
Train files: 270
Validation files: 30
Test files: 0
File type filter: .rs
Split mode: file (file-level holdout)
Each row has:
file_name: flattened source file name
text: full file contents
Training context
This dataset was used for… See the full description on the dataset page: https://huggingface.co/datasets/archit11/hyperswitch-code-corpus-track-a.code-agent-corpus
Synthetic Code-Agent Run Corpus (Projected)
⚠️ SYNTHETIC DATA — NOT REAL PRODUCTION TELEMETRY.
This dataset is synthetically generated for capacity planning / illustration of a
code-fixing agent platform's data pipeline. It does not contain real user data, real
repositories, real credentials, or real production captures. Identifiers and entity names
are anonymized (repo-A…repo-E) and run IDs/timestamps are fabricated.
What this is
A representative, labeled… See the full description on the dataset page: https://huggingface.co/datasets/Collaops/code-agent-corpus.gleam-code-corpus
Gleam Code Corpus
A structured, attributed corpus of 22,581 Gleam source files collected from 1,528 GitHub repositories for continued pre-training and code generation research.
Overview
This dataset contains .gleam source files from the Gleam programming language ecosystem. Every file includes full attribution — repo owner, name, URL, license, star count, and Hex.pm package metadata where available.
What this is for:
Continued pre-training (CPT) of code-focused… See the full description on the dataset page: https://huggingface.co/datasets/kasuboski/gleam-code-corpus.Burmese-English-Code-Mixed-Corpus
🇲🇲 Burmese-English Code-Mixed Corpus ꒰ 1,111 Rows ꒱
A high-quality, human-curated dataset of code-mixed Burmese and English sentences, specifically designed for Natural Language Processing (NLP) and Machine Learning (ML) research.
Dataset Details
Organization: DatarrX
Creator: Khant Sint Heinn (Kalix Louis)
Number of Rows: 1,111
Language: Burmese (Unicode) & English Mix
Dataset Format: .txt
License: Apache 2.0
Description
The Burmese-English Code-Mixed… See the full description on the dataset page: https://huggingface.co/datasets/hksamm/Burmese-English-Code-Mixed-Corpus.
