OpenCoder-LLM/RefineCode-code-corpus-meta
This dataset consists of meta information (including the repository name and file path) of the raw code data from RefineCode. You can collect those files referring to this metadata and reproduce RefineCode! Note: Currently, we have uploaded the meta data covered by The Stack V2 (About 50% file volume). Due to complex legal considerations, we are unable to provide the complete source code currently. We are working hard to make the remaining part available. RefineCode is a high-quality… See the full description on the dataset page: https://huggingface.co/datasets/OpenCoder-LLM/RefineCode-code-corpus-meta.
311.4k
1---2license: mit3dataset_info:4 features:5 - name: repo_name6 dtype: string7 - name: sub_path8 dtype: string9 - name: file_name10 dtype: string11 - name: file_ext12 dtype: string13 - name: file_size_in_byte14 dtype: int6415 - name: line_count16 dtype: int6417 - name: lang18 dtype: string19 - name: program_lang20 dtype: string21 - name: doc_type22 dtype: string23 splits:24 - name: The_Stack_V225 num_bytes: 4657704548526 num_examples: 33684571027 download_size: 2001908500528 dataset_size: 4657704548529configs:30- config_name: default31 data_files:32 - split: The_Stack_V233 path: data/The_Stack_V2-*34---35 36 37This dataset consists of meta information (including the repository name and file path) of the raw code data from **RefineCode**. You can collect those files referring to this metadata and reproduce **RefineCode**!38 39***Note:** Currently, we have uploaded the meta data covered by The Stack V2 (About 50% file volume). Due to complex legal considerations, we are unable to provide the complete source code currently. We are working hard to make the remaining part available.*40 41---42**RefineCode** is a **high-quality**, **reproducible** code pretraining corpus comprising **960 billion** tokens across **607** programming languages and **75 billion** code-related token recalled from web corpus, incorporating over **130** language-specific rules with customized weight assignments. 43Our dataset shows better training efficacy and efficiency compared with the training subset of The Stack V2.44 45<img src="https://raw.githubusercontent.com/OpenCoder-llm/opencoder-llm.github.io/refs/heads/main/static/images/opencoder_banner.png" alt="OpenCoder banner" style="zoom:30%;" />46 47We also use PCA to visualize the embeddings extracted from CodeBERT for The Stack V2 and **RefineCode**, showing a clear advance of our pretraining dataset.48 49<img src="https://raw.githubusercontent.com/OpenCoder-llm/opencoder-llm.github.io/refs/heads/main/static/images/compare_refinecode_stack_v2.jpg" alt="Distribution Comparsion" style="zoom:50%;" />50 51 