datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
opc-fineweb-code-corpus
OpenCoder Dataset
The OpenCoder dataset is composed of the following datasets:
opc-sft-stage1: the sft data used for opencoder sft-stage1
opc-sft-stage2: the sft data used for opencoder sft-stage2
opc-annealing-corpus: the synthetic data & algorithmic corpus used for opencoder annealing
opc-fineweb-code-corpus: the code-related page recalled from fineweb <-- you are here
opc-fineweb-math-corpus: the math-related page recalled from finewebrefineCode-code-corpus-meta: the meta-data… See the full description on the dataset page: https://huggingface.co/datasets/OpenCoder-LLM/opc-fineweb-code-corpus.swallow-code-v2
SwallowCode-v2
Resources
📑 arXiv: Read our paper for detailed methodology and results at arXiv:2505.02881.
🤗 Sister Dataset: Discover SwallowMath-v2, our companion dataset for mathematical reasoning.
💻 What is it?
SwallowCode-v1 was a high-quality Python code dataset generated through an LLM-based rewriting pipeline.
However, it had two significant limitations:
(1) it was distributed under the Llama 3.3 Community License, and
(2) its size was limited to… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-code-v2.RefineCode-code-corpus-metaThis dataset consists of meta information (including the repository name and file path) of the raw code data from RefineCode. You can collect those files referring to this metadata and reproduce RefineCode!
Note: Currently, we have uploaded the meta data covered by The Stack V2 (About 50% file volume). Due to complex legal considerations, we are unable to provide the complete source code currently. We are working hard to make the remaining part available.
RefineCode is a high-quality… See the full description on the dataset page: https://huggingface.co/datasets/OpenCoder-LLM/RefineCode-code-corpus-meta.swallow-code
SwallowCode
Notice
May 21, 2025: We have deleted ablation/exp1-the-stack-v2-train-smol-ids-python because it was flagged as potentially containing unsafe data collected from the Python subset of https://huggingface.co/datasets/bigcode/the-stack-v2-train-smol-ids. However, since this dataset can be reconstructed from the-stack-v2-train-smol-ids, there is no issue in terms of reproducibility.
May 21, 2025: ClamAV has flagged “Win.Trojan.MSShellcode-88” in… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-code.details_llm-agents__tora-code-7b-v1.0
Dataset Card for Evaluation run of llm-agents/tora-code-7b-v1.0
Dataset Summary
Dataset automatically created during the evaluation run of model llm-agents/tora-code-7b-v1.0 on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_llm-agents__tora-code-7b-v1.0.details_llm-agents__tora-code-34b-v1.0
Dataset Card for Evaluation run of llm-agents/tora-code-34b-v1.0
Dataset automatically created during the evaluation run of model llm-agents/tora-code-34b-v1.0 on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_llm-agents__tora-code-34b-v1.0.llm0to1-pt-tokenized-code
LLM0to1 사전학습 토큰화본 — 코드
10B 규모 한/영 이중언어 LLM LLM0to1-10b 를 바닥부터 학습할 때 실제로 투입된
코드 코퍼스의 토큰화본. 총 27종 / 59.1B 토큰.
왜 원문 텍스트가 아니라 토큰화본인가
이 코퍼스들은 공개 데이터셋을 스트리밍으로 받아 곧바로 토큰화했고 중간 텍스트를 보관하지 않았다.
따라서 이 .ds 파일이 학습에 들어간 데이터의 유일한 사본이다.
단점만 있는 건 아니다. 토큰화본은 학습 입력 그 자체이므로,
재토큰화 과정에서 생길 수 있는 차이 없이 학습을 그대로 재현할 수 있다.
원본 출처
bigcode/starcoderdata(언어별 서브셋) · bigcode/commitpackft · bigcode/jupyter-code-text-pairs · deepmind/code_contests
code_c 와 code_c2 처럼 2 가 붙은 것은 정제기… See the full description on the dataset page: https://huggingface.co/datasets/izlley2/llm0to1-pt-tokenized-code.details_GeorgiaTechResearchInstitute__starcoder-gpteacher-code-instruct
Dataset Card for Evaluation run of GeorgiaTechResearchInstitute/starcoder-gpteacher-code-instruct
Dataset Summary
Dataset automatically created during the evaluation run of model GeorgiaTechResearchInstitute/starcoder-gpteacher-code-instruct on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_GeorgiaTechResearchInstitute__starcoder-gpteacher-code-instruct.1gpu-llm-pretraining-corpus-15b-en-it-code
1GPU LLM Pretraining Corpus 15B EN-IT-CODE
1gpu-llm-pretraining-corpus-15b-en-it-code is the canonical document-level pretraining corpus used for the 1GPU LLM family.
It was built for training language models from scratch on a mixture of English, Italian and source code.
This Hugging Face release contains the clean, deduplicated, document-level corpus. It is intentionally published before tokenization and packing so that the training representation can be deterministically… See the full description on the dataset page: https://huggingface.co/datasets/nazdef/1gpu-llm-pretraining-corpus-15b-en-it-code.details_ajibawa-2023__OpenHermes-2.5-Code-290k-13B
Dataset Card for Evaluation run of ajibawa-2023/OpenHermes-2.5-Code-290k-13B
Dataset automatically created during the evaluation run of model ajibawa-2023/OpenHermes-2.5-Code-290k-13B on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_ajibawa-2023__OpenHermes-2.5-Code-290k-13B.big-code-2-metadataA subset of big-code-2 containing only the files:
Not in big-code-1.
Under a permissive license.
details_itsliupeng__llama2_7b_code
Dataset Card for Evaluation run of itsliupeng/llama2_7b_code
Dataset Summary
Dataset automatically created during the evaluation run of model itsliupeng/llama2_7b_code on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_itsliupeng__llama2_7b_code.details_nextai-team__Moe-3x7b-QA-Code-Inst
Dataset Card for Evaluation run of nextai-team/Moe-3x7b-QA-Code-Inst
Dataset automatically created during the evaluation run of model nextai-team/Moe-3x7b-QA-Code-Inst on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_nextai-team__Moe-3x7b-QA-Code-Inst.details_Plaban81__Moe-4x7b-math-reason-code
Dataset Card for Evaluation run of Plaban81/Moe-4x7b-math-reason-code
Dataset automatically created during the evaluation run of model Plaban81/Moe-4x7b-math-reason-code on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Plaban81__Moe-4x7b-math-reason-code.details_ajibawa-2023__Code-Mistral-7B
Dataset Card for Evaluation run of ajibawa-2023/Code-Mistral-7B
Dataset automatically created during the evaluation run of model ajibawa-2023/Code-Mistral-7B on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_ajibawa-2023__Code-Mistral-7B.details_uukuguy__speechless-code-mistral-7b-v1.0
Dataset Card for Evaluation run of uukuguy/speechless-code-mistral-7b-v1.0
Dataset Summary
Dataset automatically created during the evaluation run of model uukuguy/speechless-code-mistral-7b-v1.0 on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_uukuguy__speechless-code-mistral-7b-v1.0.details_nextai-team__Moe-2x7b-QA-Code
Dataset Card for Evaluation run of nextai-team/Moe-2x7b-QA-Code
Dataset automatically created during the evaluation run of model nextai-team/Moe-2x7b-QA-Code on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_nextai-team__Moe-2x7b-QA-Code.details_uukuguy__speechless-zephyr-code-functionary-7b
Dataset Card for Evaluation run of uukuguy/speechless-zephyr-code-functionary-7b
Dataset automatically created during the evaluation run of model uukuguy/speechless-zephyr-code-functionary-7b on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_uukuguy__speechless-zephyr-code-functionary-7b.details_ajibawa-2023__Code-290k-6.7B-Instruct
Dataset Card for Evaluation run of ajibawa-2023/Code-290k-6.7B-Instruct
Dataset automatically created during the evaluation run of model ajibawa-2023/Code-290k-6.7B-Instruct on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_ajibawa-2023__Code-290k-6.7B-Instruct.code-execution-llmllm_instruction_code_manual_yolo_lcdetails_nextai-team__Moe-4x7b-reason-code-qa
Dataset Card for Evaluation run of nextai-team/Moe-4x7b-reason-code-qa
Dataset automatically created during the evaluation run of model nextai-team/Moe-4x7b-reason-code-qa on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_nextai-team__Moe-4x7b-reason-code-qa.Josephgflowers__TinyLlama_v1.1_math_code-world-test-1-details
Dataset Card for Evaluation run of Josephgflowers/TinyLlama_v1.1_math_code-world-test-1
Dataset automatically created during the evaluation run of model Josephgflowers/TinyLlama_v1.1_math_code-world-test-1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Josephgflowers__TinyLlama_v1.1_math_code-world-test-1-details.details_ibivibiv__llama3-8b-instruct-codedetails_Undi95__Nous-Hermes-13B-Code
Dataset Card for Evaluation run of Undi95/Nous-Hermes-13B-Code
Dataset Summary
Dataset automatically created during the evaluation run of model Undi95/Nous-Hermes-13B-Code on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Undi95__Nous-Hermes-13B-Code.code-ita-dpo-small
Dataset Card for "code-instructions-ita-dpo-small"
More Information needed
details_NobodyExistsOnTheInternet__code-llama-70b-python-instruct
Dataset Card for Evaluation run of NobodyExistsOnTheInternet/code-llama-70b-python-instruct
Dataset automatically created during the evaluation run of model NobodyExistsOnTheInternet/code-llama-70b-python-instruct on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_NobodyExistsOnTheInternet__code-llama-70b-python-instruct.details_ajibawa-2023__Python-Code-33B
Dataset Card for Evaluation run of ajibawa-2023/Python-Code-33B
Dataset Summary
Dataset automatically created during the evaluation run of model ajibawa-2023/Python-Code-33B on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_ajibawa-2023__Python-Code-33B.details_layoric__llama-2-13b-code-alpaca
Dataset Card for Evaluation run of layoric/llama-2-13b-code-alpaca
Dataset Summary
Dataset automatically created during the evaluation run of model layoric/llama-2-13b-code-alpaca on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_layoric__llama-2-13b-code-alpaca.llm_instruction_code_V6.1
