Team Ai
6 results

code-corpus

OpenCoder-LLM /opc-fineweb-code-corpus OpenCoder Dataset The OpenCoder dataset is composed of the following datasets: opc-sft-stage1: the sft data used for opencoder sft-stage1 opc-sft-stage2: the sft data used for opencoder sft-stage2 opc-annealing-corpus: the synthetic data & algorithmic corpus used for opencoder annealing opc-fineweb-code-corpus: the code-related page recalled from fineweb <-- you are here opc-fineweb-math-corpus: the math-related page recalled from finewebrefineCode-code-corpus-meta: the meta-data… See the full description on the dataset page: https://huggingface.co/datasets/OpenCoder-LLM/opc-fineweb-code-corpus.tabular100M<n<1B64 likes7.2k downloads2y agoHugging Facerufatronics /clean-code-corpus-v1text1K<n<10K0 likes1.7k downloads7h agoHugging FaceOpenCoder-LLM /RefineCode-code-corpus-metaThis dataset consists of meta information (including the repository name and file path) of the raw code data from RefineCode. You can collect those files referring to this metadata and reproduce RefineCode! Note: Currently, we have uploaded the meta data covered by The Stack V2 (About 50% file volume). Due to complex legal considerations, we are unable to provide the complete source code currently. We are working hard to make the remaining part available. RefineCode is a high-quality… See the full description on the dataset page: https://huggingface.co/datasets/OpenCoder-LLM/RefineCode-code-corpus-meta.tabular100M<n<1B30 likes1.4k downloads2y agoHugging FaceCorpus-NZ /High-Quality-Code High-Quality-Code: Synthetic + Real (MAXIMUM CODE) A massive, high-quality code dataset built with maximum code philosophy – as much code as possible. Components Synthetic syntax-correction dataset – 5M+ examples across 33 languages (original code_syntax_dataset_1GB.csv) Real high-quality code from GitHub – 500 top-starred repositories – BOTH zips and extracted source Current Status: IN PROGRESS Target: 500 repos Currently uploaded: 12 extracted… See the full description on the dataset page: https://huggingface.co/datasets/Corpus-NZ/High-Quality-Code.100M<n<1B0 likes1.3k downloads27d agoHugging Faceteven /code_docstring_corpusHF version of Edinburgh-NLP's Code docstrings corpus text100K<n<1M8 likes319 downloads4y agoHugging Facenazdef /1gpu-llm-pretraining-corpus-15b-en-it-code 1GPU LLM Pretraining Corpus 15B EN-IT-CODE 1gpu-llm-pretraining-corpus-15b-en-it-code is the canonical document-level pretraining corpus used for the 1GPU LLM family. It was built for training language models from scratch on a mixture of English, Italian and source code. This Hugging Face release contains the clean, deduplicated, document-level corpus. It is intentionally published before tokenization and packing so that the training representation can be deterministically… See the full description on the dataset page: https://huggingface.co/datasets/nazdef/1gpu-llm-pretraining-corpus-15b-en-it-code.text-generation0 likes318 downloads11d agoHugging Face