Team Ai
Datasetpublic

OpenCoder-LLM/RefineCode-code-corpus-meta

This dataset consists of meta information (including the repository name and file path) of the raw code data from RefineCode. You can collect those files referring to this metadata and reproduce RefineCode! Note: Currently, we have uploaded the meta data covered by The Stack V2 (About 50% file volume). Due to complex legal considerations, we are unable to provide the complete source code currently. We are working hard to make the remaining part available. RefineCode is a high-quality… See the full description on the dataset page: https://huggingface.co/datasets/OpenCoder-LLM/RefineCode-code-corpus-meta.

sourceHugging Facemitupdated 2y agoView on Hugging Face
31likes1.4kdownloads
README.md51 linesDownload Raw Back to root
1---2license: mit3dataset_info:4  features:5  - name: repo_name6    dtype: string7  - name: sub_path8    dtype: string9  - name: file_name10    dtype: string11  - name: file_ext12    dtype: string13  - name: file_size_in_byte14    dtype: int6415  - name: line_count16    dtype: int6417  - name: lang18    dtype: string19  - name: program_lang20    dtype: string21  - name: doc_type22    dtype: string23  splits:24  - name: The_Stack_V225    num_bytes: 4657704548526    num_examples: 33684571027  download_size: 2001908500528  dataset_size: 4657704548529configs:30- config_name: default31  data_files:32  - split: The_Stack_V233    path: data/The_Stack_V2-*34---35 36 37This dataset consists of meta information (including the repository name and file path) of the raw code data from **RefineCode**. You can collect those files referring to this metadata and reproduce **RefineCode**!38 39***Note:** Currently, we have uploaded the meta data covered by The Stack V2 (About 50% file volume). Due to complex legal considerations, we are unable to provide the complete source code currently. We are working hard to make the remaining part available.*40 41---42**RefineCode** is a **high-quality**, **reproducible** code pretraining corpus comprising **960 billion** tokens across **607** programming languages and **75 billion** code-related token recalled from web corpus, incorporating over **130** language-specific rules with customized weight assignments. 43Our dataset shows better training efficacy and efficiency compared with the training subset of The Stack V2.44 45<img src="https://raw.githubusercontent.com/OpenCoder-llm/opencoder-llm.github.io/refs/heads/main/static/images/opencoder_banner.png" alt="OpenCoder banner" style="zoom:30%;" />46 47We also use PCA to visualize the embeddings extracted from CodeBERT for The Stack V2 and **RefineCode**, showing a clear advance of our pretraining dataset.48 49<img src="https://raw.githubusercontent.com/OpenCoder-llm/opencoder-llm.github.io/refs/heads/main/static/images/compare_refinecode_stack_v2.jpg" alt="Distribution Comparsion" style="zoom:50%;" />50 51