Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01OpenCoder-LLM /opc-fineweb-code-corpus OpenCoder Dataset The OpenCoder dataset is composed of the following datasets: opc-sft-stage1: the sft data used for opencoder sft-stage1 opc-sft-stage2: the sft data used for opencoder sft-stage2 opc-annealing-corpus: the synthetic data & algorithmic corpus used for opencoder annealing opc-fineweb-code-corpus: the code-related page recalled from fineweb <-- you are here opc-fineweb-math-corpus: the math-related page recalled from finewebrefineCode-code-corpus-meta: the meta-data… See the full description on the dataset page: https://huggingface.co/datasets/OpenCoder-LLM/opc-fineweb-code-corpus.tabular100M<n<1B65 likes7.2k downloads2y agoHugging Face02tokyotech-llm /swallow-code-v2 SwallowCode-v2 Resources 📑 arXiv: Read our paper for detailed methodology and results at arXiv:2505.02881. 🤗 Sister Dataset: Discover SwallowMath-v2, our companion dataset for mathematical reasoning. 💻 What is it? SwallowCode-v1 was a high-quality Python code dataset generated through an LLM-based rewriting pipeline. However, it had two significant limitations: (1) it was distributed under the Llama 3.3 Community License, and (2) its size was limited to… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-code-v2.tabulartext-generation100M<n<1B53 likes5k downloads11mo agoHugging Face03OpenCoder-LLM /RefineCode-code-corpus-metaThis dataset consists of meta information (including the repository name and file path) of the raw code data from RefineCode. You can collect those files referring to this metadata and reproduce RefineCode! Note: Currently, we have uploaded the meta data covered by The Stack V2 (About 50% file volume). Due to complex legal considerations, we are unable to provide the complete source code currently. We are working hard to make the remaining part available. RefineCode is a high-quality… See the full description on the dataset page: https://huggingface.co/datasets/OpenCoder-LLM/RefineCode-code-corpus-meta.tabular100M<n<1B31 likes1.4k downloads2y agoHugging Face04tokyotech-llm /swallow-code SwallowCode Notice May 21, 2025: We have deleted ablation/exp1-the-stack-v2-train-smol-ids-python because it was flagged as potentially containing unsafe data collected from the Python subset of https://huggingface.co/datasets/bigcode/the-stack-v2-train-smol-ids. However, since this dataset can be reconstructed from the-stack-v2-train-smol-ids, there is no issue in terms of reproducibility. May 21, 2025: ClamAV has flagged “Win.Trojan.MSShellcode-88” in… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-code.tabulartext-generation100M<n<1B76 likes868 downloads7mo agoHugging Face05open-llm-leaderboard-old /details_llm-agents__tora-code-7b-v1.0 Dataset Card for Evaluation run of llm-agents/tora-code-7b-v1.0 Dataset Summary Dataset automatically created during the evaluation run of model llm-agents/tora-code-7b-v1.0 on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_llm-agents__tora-code-7b-v1.0.5 likes757 downloads3y agoHugging Face06open-llm-leaderboard-old /details_llm-agents__tora-code-34b-v1.0 Dataset Card for Evaluation run of llm-agents/tora-code-34b-v1.0 Dataset automatically created during the evaluation run of model llm-agents/tora-code-34b-v1.0 on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_llm-agents__tora-code-34b-v1.0.1 likes676 downloads3y agoHugging Face07izlley2 /llm0to1-pt-tokenized-code LLM0to1 사전학습 토큰화본 — 코드 10B 규모 한/영 이중언어 LLM LLM0to1-10b 를 바닥부터 학습할 때 실제로 투입된 코드 코퍼스의 토큰화본. 총 27종 / 59.1B 토큰. 왜 원문 텍스트가 아니라 토큰화본인가 이 코퍼스들은 공개 데이터셋을 스트리밍으로 받아 곧바로 토큰화했고 중간 텍스트를 보관하지 않았다. 따라서 이 .ds 파일이 학습에 들어간 데이터의 유일한 사본이다. 단점만 있는 건 아니다. 토큰화본은 학습 입력 그 자체이므로, 재토큰화 과정에서 생길 수 있는 차이 없이 학습을 그대로 재현할 수 있다. 원본 출처 bigcode/starcoderdata(언어별 서브셋) · bigcode/commitpackft · bigcode/jupyter-code-text-pairs · deepmind/code_contests code_c 와 code_c2 처럼 2 가 붙은 것은 정제기… See the full description on the dataset page: https://huggingface.co/datasets/izlley2/llm0to1-pt-tokenized-code.text-generationn>1T0 likes382 downloads2mo agoHugging Face08open-llm-leaderboard-old /details_GeorgiaTechResearchInstitute__starcoder-gpteacher-code-instruct Dataset Card for Evaluation run of GeorgiaTechResearchInstitute/starcoder-gpteacher-code-instruct Dataset Summary Dataset automatically created during the evaluation run of model GeorgiaTechResearchInstitute/starcoder-gpteacher-code-instruct on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_GeorgiaTechResearchInstitute__starcoder-gpteacher-code-instruct.4 likes329 downloads3y agoHugging Face09nazdef /1gpu-llm-pretraining-corpus-15b-en-it-code 1GPU LLM Pretraining Corpus 15B EN-IT-CODE 1gpu-llm-pretraining-corpus-15b-en-it-code is the canonical document-level pretraining corpus used for the 1GPU LLM family. It was built for training language models from scratch on a mixture of English, Italian and source code. This Hugging Face release contains the clean, deduplicated, document-level corpus. It is intentionally published before tokenization and packing so that the training representation can be deterministically… See the full description on the dataset page: https://huggingface.co/datasets/nazdef/1gpu-llm-pretraining-corpus-15b-en-it-code.text-generation0 likes318 downloads12d agoHugging Face10open-llm-leaderboard-old /details_ajibawa-2023__OpenHermes-2.5-Code-290k-13B Dataset Card for Evaluation run of ajibawa-2023/OpenHermes-2.5-Code-290k-13B Dataset automatically created during the evaluation run of model ajibawa-2023/OpenHermes-2.5-Code-290k-13B on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_ajibawa-2023__OpenHermes-2.5-Code-290k-13B.0 likes294 downloads3y agoHugging Face11LLMDH /big-code-2-metadataA subset of big-code-2 containing only the files: Not in big-code-1. Under a permissive license. text100M<n<1B0 likes249 downloads2y agoHugging Face12open-llm-leaderboard-old /details_itsliupeng__llama2_7b_code Dataset Card for Evaluation run of itsliupeng/llama2_7b_code Dataset Summary Dataset automatically created during the evaluation run of model itsliupeng/llama2_7b_code on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_itsliupeng__llama2_7b_code.0 likes200 downloads3y agoHugging Face13open-llm-leaderboard-old /details_nextai-team__Moe-3x7b-QA-Code-Inst Dataset Card for Evaluation run of nextai-team/Moe-3x7b-QA-Code-Inst Dataset automatically created during the evaluation run of model nextai-team/Moe-3x7b-QA-Code-Inst on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_nextai-team__Moe-3x7b-QA-Code-Inst.0 likes197 downloads3y agoHugging Face14open-llm-leaderboard-old /details_Plaban81__Moe-4x7b-math-reason-code Dataset Card for Evaluation run of Plaban81/Moe-4x7b-math-reason-code Dataset automatically created during the evaluation run of model Plaban81/Moe-4x7b-math-reason-code on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Plaban81__Moe-4x7b-math-reason-code.0 likes159 downloads3y agoHugging Face15open-llm-leaderboard-old /details_ajibawa-2023__Code-Mistral-7B Dataset Card for Evaluation run of ajibawa-2023/Code-Mistral-7B Dataset automatically created during the evaluation run of model ajibawa-2023/Code-Mistral-7B on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_ajibawa-2023__Code-Mistral-7B.0 likes147 downloads3y agoHugging Face16open-llm-leaderboard-old /details_uukuguy__speechless-code-mistral-7b-v1.0 Dataset Card for Evaluation run of uukuguy/speechless-code-mistral-7b-v1.0 Dataset Summary Dataset automatically created during the evaluation run of model uukuguy/speechless-code-mistral-7b-v1.0 on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_uukuguy__speechless-code-mistral-7b-v1.0.0 likes141 downloads3y agoHugging Face17open-llm-leaderboard-old /details_nextai-team__Moe-2x7b-QA-Code Dataset Card for Evaluation run of nextai-team/Moe-2x7b-QA-Code Dataset automatically created during the evaluation run of model nextai-team/Moe-2x7b-QA-Code on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_nextai-team__Moe-2x7b-QA-Code.0 likes137 downloads3y agoHugging Face18open-llm-leaderboard-old /details_uukuguy__speechless-zephyr-code-functionary-7b Dataset Card for Evaluation run of uukuguy/speechless-zephyr-code-functionary-7b Dataset automatically created during the evaluation run of model uukuguy/speechless-zephyr-code-functionary-7b on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_uukuguy__speechless-zephyr-code-functionary-7b.0 likes129 downloads3y agoHugging Face19open-llm-leaderboard-old /details_ajibawa-2023__Code-290k-6.7B-Instruct Dataset Card for Evaluation run of ajibawa-2023/Code-290k-6.7B-Instruct Dataset automatically created during the evaluation run of model ajibawa-2023/Code-290k-6.7B-Instruct on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_ajibawa-2023__Code-290k-6.7B-Instruct.0 likes127 downloads3y agoHugging Face20Guilherme34 /code-execution-llmtabular1K<n<10K2 likes123 downloads2y agoHugging Face21Shiveswarran /llm_instruction_code_manual_yolo_lctextn<1K4 likes108 downloads3y agoHugging Face22open-llm-leaderboard-old /details_nextai-team__Moe-4x7b-reason-code-qa Dataset Card for Evaluation run of nextai-team/Moe-4x7b-reason-code-qa Dataset automatically created during the evaluation run of model nextai-team/Moe-4x7b-reason-code-qa on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_nextai-team__Moe-4x7b-reason-code-qa.0 likes90 downloads3y agoHugging Face23open-llm-leaderboard /Josephgflowers__TinyLlama_v1.1_math_code-world-test-1-detailsgated Dataset Card for Evaluation run of Josephgflowers/TinyLlama_v1.1_math_code-world-test-1 Dataset automatically created during the evaluation run of model Josephgflowers/TinyLlama_v1.1_math_code-world-test-1 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Josephgflowers__TinyLlama_v1.1_math_code-world-test-1-details.tabular10K<n<100K0 likes87 downloads2y agoHugging Face24open-llm-leaderboard-old /details_ibivibiv__llama3-8b-instruct-code0 likes85 downloads2y agoHugging Face25open-llm-leaderboard-old /details_Undi95__Nous-Hermes-13B-Code Dataset Card for Evaluation run of Undi95/Nous-Hermes-13B-Code Dataset Summary Dataset automatically created during the evaluation run of model Undi95/Nous-Hermes-13B-Code on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Undi95__Nous-Hermes-13B-Code.0 likes82 downloads3y agoHugging Face26mii-llm /code-ita-dpo-small Dataset Card for "code-instructions-ita-dpo-small" More Information needed textn<1K3 likes76 downloads3y agoHugging Face27open-llm-leaderboard-old /details_NobodyExistsOnTheInternet__code-llama-70b-python-instruct Dataset Card for Evaluation run of NobodyExistsOnTheInternet/code-llama-70b-python-instruct Dataset automatically created during the evaluation run of model NobodyExistsOnTheInternet/code-llama-70b-python-instruct on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_NobodyExistsOnTheInternet__code-llama-70b-python-instruct.0 likes73 downloads3y agoHugging Face28open-llm-leaderboard-old /details_ajibawa-2023__Python-Code-33B Dataset Card for Evaluation run of ajibawa-2023/Python-Code-33B Dataset Summary Dataset automatically created during the evaluation run of model ajibawa-2023/Python-Code-33B on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_ajibawa-2023__Python-Code-33B.0 likes72 downloads3y agoHugging Face29open-llm-leaderboard-old /details_layoric__llama-2-13b-code-alpaca Dataset Card for Evaluation run of layoric/llama-2-13b-code-alpaca Dataset Summary Dataset automatically created during the evaluation run of model layoric/llama-2-13b-code-alpaca on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_layoric__llama-2-13b-code-alpaca.0 likes71 downloads3y agoHugging Face30Shiveswarran /llm_instruction_code_V6.1text100K<n<1M7 likes71 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.