Team Ai
15 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01OpenCoder-LLM /opc-fineweb-code-corpus OpenCoder Dataset The OpenCoder dataset is composed of the following datasets: opc-sft-stage1: the sft data used for opencoder sft-stage1 opc-sft-stage2: the sft data used for opencoder sft-stage2 opc-annealing-corpus: the synthetic data & algorithmic corpus used for opencoder annealing opc-fineweb-code-corpus: the code-related page recalled from fineweb <-- you are here opc-fineweb-math-corpus: the math-related page recalled from finewebrefineCode-code-corpus-meta: the meta-data… See the full description on the dataset page: https://huggingface.co/datasets/OpenCoder-LLM/opc-fineweb-code-corpus.tabular100M<n<1B66 likes7k downloads2y agoHugging Face02rufatronics /clean-code-corpus-v1text1K<n<10K0 likes3.1k downloads1h agoHugging Face03OpenCoder-LLM /RefineCode-code-corpus-metaThis dataset consists of meta information (including the repository name and file path) of the raw code data from RefineCode. You can collect those files referring to this metadata and reproduce RefineCode! Note: Currently, we have uploaded the meta data covered by The Stack V2 (About 50% file volume). Due to complex legal considerations, we are unable to provide the complete source code currently. We are working hard to make the remaining part available. RefineCode is a high-quality… See the full description on the dataset page: https://huggingface.co/datasets/OpenCoder-LLM/RefineCode-code-corpus-meta.tabular100M<n<1B32 likes1.3k downloads2y agoHugging Face04teven /code_docstring_corpusHF version of Edinburgh-NLP's Code docstrings corpus text100K<n<1M8 likes299 downloads4y agoHugging Face05archit11 /verl-code-corpus-track-a-file-split archit11/verl-code-corpus-track-a-file-split Repository-specific code corpus extracted from the verl project and split by file for training/evaluation. What is in this dataset Source corpus: data/code_corpus_verl Total files: 214 Train files: 172 Validation files: 21 Test files: 21 File type filter: .py Split mode: file (file-level holdout) Each row has: file_name: flattened source file name text: full file contents Training context This dataset was used… See the full description on the dataset page: https://huggingface.co/datasets/archit11/verl-code-corpus-track-a-file-split.texttext-generationn<1K0 likes108 downloads8mo agoHugging Face06burtenshaw /PleIAs_common_corpus_code_classificationtext100K<n<1M1 likes95 downloads1y agoHugging Face07HuggingPanda /SDAIANCAI-Saudilang-Code-Switch-Corpusaudio1K<n<10K2 likes56 downloads2y agoHugging Face08krystv /code-corpus-llm-training Code Corpus for LLM Training Manually collected from top open-source repositories across: video/graphics editors, browsers, terminals, UI/UX, Qt/QML, Flutter, Rust, Python, ethical hacking, system-level, game engines, web frameworks, and more. Stats Records: 240,378 Raw text: 2,156,908,643 chars (~2.01 GB) Domains: 20 Domains web_ui: 32,354 records cpp: 29,792 records kotlin_android: 19,476 records ui_ux_design: 19,382 records rust: 15,440 records python:… See the full description on the dataset page: https://huggingface.co/datasets/krystv/code-corpus-llm-training.text100K<n<1M0 likes47 downloads5mo agoHugging Face09archit11 /hyperswitch-code-corpus-track-a archit11/hyperswitch-code-corpus-track-a Repository-specific code corpus extracted from hyperswitch and split by file for training/evaluation. What is in this dataset Source corpus: data/code_corpus_hyperswitch Total files: 300 Train files: 270 Validation files: 30 Test files: 0 File type filter: .rs Split mode: file (file-level holdout) Each row has: file_name: flattened source file name text: full file contents Training context This dataset was used for… See the full description on the dataset page: https://huggingface.co/datasets/archit11/hyperswitch-code-corpus-track-a.texttext-generationn<1K0 likes41 downloads8mo agoHugging Face10seongchaeae /sage-code-corpus-v1text1M<n<10M1 likes36 downloads5mo agoHugging Face11archit11 /verl-code-corpustextn<1K0 likes28 downloads8mo agoHugging Face12kasuboski /gleam-code-corpus Gleam Code Corpus A structured, attributed corpus of 22,581 Gleam source files collected from 1,528 GitHub repositories for continued pre-training and code generation research. Overview This dataset contains .gleam source files from the Gleam programming language ecosystem. Every file includes full attribution — repo owner, name, URL, license, star count, and Hex.pm package metadata where available. What this is for: Continued pre-training (CPT) of code-focused… See the full description on the dataset page: https://huggingface.co/datasets/kasuboski/gleam-code-corpus.tabulartext-generation10K<n<100K0 likes28 downloads5mo agoHugging Face13coding-assistant-custom /mini-code-corpus Dataset Card for "mini-code-corpus" More Information needed textn<1K1 likes27 downloads3y agoHugging Face14ostapeno /opc-annealing-corpus-synth-qa-code_python_js_tstext1M<n<10M0 likes22 downloads2y agoHugging Face15Wodous /Code_corpustext100K<n<1M1 likes6 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.