Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01OpenCoder-LLM /opc-fineweb-code-corpus OpenCoder Dataset The OpenCoder dataset is composed of the following datasets: opc-sft-stage1: the sft data used for opencoder sft-stage1 opc-sft-stage2: the sft data used for opencoder sft-stage2 opc-annealing-corpus: the synthetic data & algorithmic corpus used for opencoder annealing opc-fineweb-code-corpus: the code-related page recalled from fineweb <-- you are here opc-fineweb-math-corpus: the math-related page recalled from finewebrefineCode-code-corpus-meta: the meta-data… See the full description on the dataset page: https://huggingface.co/datasets/OpenCoder-LLM/opc-fineweb-code-corpus.tabular100M<n<1B66 likes7k downloads2y agoHugging Face02OpenCoder-LLM /RefineCode-code-corpus-metaThis dataset consists of meta information (including the repository name and file path) of the raw code data from RefineCode. You can collect those files referring to this metadata and reproduce RefineCode! Note: Currently, we have uploaded the meta data covered by The Stack V2 (About 50% file volume). Due to complex legal considerations, we are unable to provide the complete source code currently. We are working hard to make the remaining part available. RefineCode is a high-quality… See the full description on the dataset page: https://huggingface.co/datasets/OpenCoder-LLM/RefineCode-code-corpus-meta.tabular100M<n<1B32 likes1.3k downloads2y agoHugging Face03LLMDH /big-code-2-metadataA subset of big-code-2 containing only the files: Not in big-code-1. Under a permissive license. text100M<n<1B0 likes220 downloads2y agoHugging Face04Guilherme34 /code-execution-llmtabular1K<n<10K2 likes123 downloads2y agoHugging Face05mii-llm /code-ita-dpo-small Dataset Card for "code-instructions-ita-dpo-small" More Information needed textn<1K4 likes92 downloads3y agoHugging Face06rubricreward /llm-metric-ace-code-pairwisetext100K<n<1M0 likes48 downloads1y agoHugging Face07krystv /code-corpus-llm-training Code Corpus for LLM Training Manually collected from top open-source repositories across: video/graphics editors, browsers, terminals, UI/UX, Qt/QML, Flutter, Rust, Python, ethical hacking, system-level, game engines, web frameworks, and more. Stats Records: 240,378 Raw text: 2,156,908,643 chars (~2.01 GB) Domains: 20 Domains web_ui: 32,354 records cpp: 29,792 records kotlin_android: 19,476 records ui_ux_design: 19,382 records rust: 15,440 records python:… See the full description on the dataset page: https://huggingface.co/datasets/krystv/code-corpus-llm-training.text100K<n<1M0 likes47 downloads5mo agoHugging Face08sungjin-code /amazon-reviews-for-llm-extended Cross-domain sequential recommendation dataset A sequential recommendation dataset drawn from Amazon Reviews 2023, covering Books, CDs_and_Vinyl, Movies_and_TV, Video_Games. Each row of interactions.parquet is one user buying or reviewing one item at one time. Users are sampled so that every one of them is active in all domains, their interactions are ordered chronologically and cut into train/valid/test, and each interaction carries a fixed set of 10 candidate items for ranking… See the full description on the dataset page: https://huggingface.co/datasets/sungjin-code/amazon-reviews-for-llm-extended.tabular1K<n<10K0 likes45 downloads2mo agoHugging Face09sungjin-code /amazon-reviews-books-for-llm Books sequential recommendation dataset A sequential recommendation dataset drawn from Amazon Reviews 2023, covering Books. Each row of interactions.parquet is one user buying or reviewing one item at one time. Users are sampled so that every one of them is active in all domains, their interactions are ordered chronologically and cut into train/valid/test, and each interaction carries a fixed set of 10 candidate items for ranking evaluation. Integer user_idx / item_idx columns… See the full description on the dataset page: https://huggingface.co/datasets/sungjin-code/amazon-reviews-books-for-llm.tabular10K<n<100K0 likes38 downloads2mo agoHugging Face10sungjin-code /amazon-reviews-for-llm Cross-domain sequential recommendation dataset A sequential recommendation dataset drawn from Amazon Reviews 2023, covering Books, CDs_and_Vinyl, Movies_and_TV, Video_Games. Each row of interactions.parquet is one user buying or reviewing one item at one time. Users are sampled so that every one of them is active in all domains, their interactions are ordered chronologically and cut into train/valid/test, and each interaction carries a fixed set of 10 candidate items for ranking… See the full description on the dataset page: https://huggingface.co/datasets/sungjin-code/amazon-reviews-for-llm.tabular1K<n<10K0 likes36 downloads2mo agoHugging Face11bermaneh /pde-llm-eval-code-perturbation-dataset pde-llm-eval-code-perturbation-dataset Code Perturbation Dataset (final). 256 PDE solver implementations: 64 base programs (32 physically valid, 32 with an injected bug that still runs to completion) expanded by 4 lexical perturbation conditions each -- original comments, comments removed, comments swapped in from a different implementation, and descriptive identifiers replaced by meaningless placeholders. The perturbations change the lexical surface only; executable behaviour… See the full description on the dataset page: https://huggingface.co/datasets/bermaneh/pde-llm-eval-code-perturbation-dataset.tabularn<1K0 likes34 downloads1mo agoHugging Face12rubricreward /llm-metric-ace-code-pairwise-newtext100K<n<1M0 likes33 downloads1y agoHugging Face13DCAgent /llm-verifier-code-conteststextn<1K0 likes20 downloads10mo agoHugging Face14benjamintli /code-retrieval-combined-v2-llm-negativestext1M<n<10M0 likes13 downloads7mo agoHugging Face15benjamintli /code-retrieval-hard-negatives-llm-verified-mergedtext100K<n<1M0 likes11 downloads7mo agoHugging Face16DCAgent /llm-verifier-code-contests-noblocktextn<1K0 likes10 downloads7mo agoHugging Face17mlfoundations-dev /d1_code_mc_llm_eval_636d mlfoundations-dev/d1_code_mc_llm_eval_636d Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces Accuracy 23.0 63.3 74.6 27.8 44.9 45.1 45.5 15.3 18.9 AIME24 Average Accuracy: 23.00% ± 1.45% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1 23.33% 7 30 2 23.33% 7 30 3 23.33% 7 30 4 20.00% 6 30… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/d1_code_mc_llm_eval_636d.tabular1K<n<10K0 likes8 downloads1y agoHugging Face18mlfoundations-dev /d1_code_mc_llm_1k_eval_636d mlfoundations-dev/d1_code_mc_llm_1k_eval_636d Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces Accuracy 18.7 53.8 73.2 28.8 40.3 38.9 25.7 6.4 8.6 AIME24 Average Accuracy: 18.67% ± 1.17% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1 20.00% 6 30 2 20.00% 6 30 3 20.00% 6 30 4 26.67% 8 30… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/d1_code_mc_llm_1k_eval_636d.tabular1K<n<10K0 likes8 downloads1y agoHugging Face19mlfoundations-dev /d1_code_mc_llmtabular10K<n<100K0 likes7 downloads1y agoHugging Face20mlfoundations-dev /d1_code_mc_llm_3k_eval_636d mlfoundations-dev/d1_code_mc_llm_3k_eval_636d Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces Accuracy 23.3 61.0 75.6 28.4 44.5 42.8 33.1 9.5 12.4 AIME24 Average Accuracy: 23.33% ± 1.49% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1 26.67% 8 30 2 30.00% 9 30 3 20.00% 6 30 4 23.33% 7 30… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/d1_code_mc_llm_3k_eval_636d.tabular1K<n<10K0 likes7 downloads1y agoHugging Face21mlfoundations-dev /d1_code_mc_llm_10k_eval_636d mlfoundations-dev/d1_code_mc_llm_10k_eval_636d Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces Accuracy 19.7 63.5 77.2 28.4 47.0 42.1 38.7 10.5 15.2 AIME24 Average Accuracy: 19.67% ± 1.45% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1 16.67% 5 30 2 13.33% 4 30 3 23.33% 7 30 4 26.67% 8… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/d1_code_mc_llm_10k_eval_636d.tabular1K<n<10K0 likes7 downloads1y agoHugging Face22beimnet777 /LLM_generated_Amharic_QA_for_Family_Code_of_Ethiopiagatedtext1K<n<10K0 likes6 downloads2y agoHugging Face23kimdsy /code_llm2textn<1K0 likes6 downloads2y agoHugging Face24rubricreward /llm-metric-code-contest-pairwisegatedtext10K<n<100K0 likes6 downloads1y agoHugging Face25mlfoundations-dev /d1_code_mc_llm_3ktabular1K<n<10K0 likes6 downloads1y agoHugging Face26mlfoundations-dev /d1_code_mc_llm_10ktabular1K<n<10K0 likes6 downloads1y agoHugging Face27mlfoundations-dev /d1_code_mc_llm_0.3k_eval_636d mlfoundations-dev/d1_code_mc_llm_0.3k_eval_636d Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces Accuracy 16.0 55.2 71.8 27.2 39.1 37.5 26.8 6.6 6.6 AIME24 Average Accuracy: 16.00% ± 1.40% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1 26.67% 8 30 2 16.67% 5 30 3 13.33% 4 30 4 16.67% 5 30… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/d1_code_mc_llm_0.3k_eval_636d.tabular1K<n<10K0 likes6 downloads1y agoHugging Face28llm-crafter /l4-08-code-generation-datatextn<1K0 likes6 downloads8mo agoHugging Face29asishley /LLMSecEval-Prompts_generated_codetextn<1K0 likes6 downloads7mo agoHugging Face30mlfoundations-dev /d2_code_mc_llmtabularn<1K0 likes5 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.