Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01kaysss /leetcode-problem-solutions LeetCode Solution Dataset This dataset contains community-contributed LeetCode solutions scraped from public discussions and solution pages, enriched with metadata such as vote counts, author info, tags, and full code content. The goal is to make high-quality, peer-reviewed coding solutions programmatically accessible for research, analysis, educational use, or developer tooling. Column Descriptions Column Name Type Description question_slug string The unique… See the full description on the dataset page: https://huggingface.co/datasets/kaysss/leetcode-problem-solutions.tabulartext-classification100K<n<1M9 likes5.4k downloads1y agoHugging Face02davidheineman /deepseek-leetcodeDeepseek Leetcode dataset from https://github.com/deepseek-ai/DeepSeek-Coder/tree/main/Evaluation/LeetCode tabularn<1K0 likes850 downloads1y agoHugging Face03kaysss /leetcode-problem-set LeetCode Scraper Dataset This dataset contains information scraped from LeetCode. It is designed to assist developers in analyzing LeetCode problems, generating insights, and building tools for competitive programming or educational purposes. Dataset Contents The dataset includes the following files: problem_set.csv Contains a list of LeetCode problems with metadata such as difficulty, acceptance rate, tags, and more. Columns: acRate: Acceptance rate of the… See the full description on the dataset page: https://huggingface.co/datasets/kaysss/leetcode-problem-set.tabularquestion-answering1K<n<10K10 likes804 downloads1y agoHugging Face04olegshulyakov /doocs-leetcode-solutions Doocs LeetCode Solutions LeetCode problems with solutions in many programming languages, built from the Doocs LeetCode repository. Each solution comes with the approach name, the reasoning that leads to it, and an explanation with complexity analysis. The dataset is meant for fine-tuning and evaluating code generation models. The dataset is regenerated monthly from the latest Doocs commit by the leetcode-dataset-generator tool. Structure The dataset has two… See the full description on the dataset page: https://huggingface.co/datasets/olegshulyakov/doocs-leetcode-solutions.tabulartext-generation10K<n<100K3 likes756 downloads1d agoHugging Face05kaysss /leetcode-problem-detailed LeetCode Scraper Dataset This dataset contains information scraped from LeetCode, including problem details, metadata, and related files. It is designed to assist developers in analyzing LeetCode problems, generating insights, and building tools for competitive programming or educational purposes. questions_deets.csv Contains detailed information about each problem, including problem descriptions, constraints, and examples. Columns: questionFrontendId: Unique problem ID.… See the full description on the dataset page: https://huggingface.co/datasets/kaysss/leetcode-problem-detailed.tabulartext-classification1K<n<10K11 likes720 downloads1y agoHugging Face06hi-todayis-jh /grpo-qwen3-1.7b-nemotron-leetcode-clean-3.2k-bs32-n8-verl091-epoch2-146102-rollouts Coding RL rollouts grpo_Qwen3-1.7B_Nemotron-LeetCode-clean-3.2k_bs32_n8_seqs16_32k_epoch2_verl091 One verified gzip JSONL shard per training step; 256 responses per shard. LCB binary grading after thinking, without an EOS gate. tabular10K<n<100K0 likes636 downloads8d agoHugging Face07hi-todayis-jh /grpo-qwen3-1.7b-nemotron-leetcode-clean-3.2k-bs32-n8-verl091-146102-rollouts Coding GRPO rollouts grpo_Qwen3-1.7B_Nemotron-LeetCode-clean-3.2k_bs32_n8_seqs16_32k_1epoch_verl091 One verified gzip JSONL shard per training step; 256 responses per shard. LCB binary grading after thinking, without an EOS gate. tabular10K<n<100K0 likes455 downloads9d agoHugging Face08whiskwhite /leetcode-complete Complete LeetCode Problems Dataset This dataset contains a comprehensive collection of LeetCode problems (including premium) with AI-generated solutions in JSONL format. It is regularly updated to include new problems as they are added to LeetCode. Splits The dataset is divided into the following splits: train: Contains approximately 80% of the problems for training validation: Contains approximately 10% of the problems for validation test: Contains approximately… See the full description on the dataset page: https://huggingface.co/datasets/whiskwhite/leetcode-complete.tabulartext-generation1K<n<10K1 likes381 downloads16h agoHugging Face09skandermoalla /qrpo-paper-llama-nosft-leetcode-sandbox-temp1-ref50-offpolicy10random-sandbox qrpo-paper-llama-nosft-leetcode-sandbox-temp1-ref50-offpolicy10random-sandbox Dataset with reference completions and rewards for a specific model and reward model, ready for training with the QRPO reference codebase (https://github.com/CLAIRE-Labo/quantile-reward-policy-optimization). Part of the dataset collection for the paper Quantile Reward Policy Optimization: Alignment with Pointwise Regression and Exact Partition Functions (https://arxiv.org/pdf/2507.08068). tabular10K<n<100K0 likes355 downloads10mo agoHugging Face10cassanof /leetcode-solutionsFrom: https://www.kaggle.com/datasets/jacobhds/leetcode-solutions-and-content-kpis tabular10K<n<100K11 likes244 downloads3y agoHugging Face11skandermoalla /qrpo-paper-llama-sft-leetcode-sandbox-temp1-ref50-offpolicy10random-sandbox qrpo-paper-llama-sft-leetcode-sandbox-temp1-ref50-offpolicy10random-sandbox Dataset with reference completions and rewards for a specific model and reward model, ready for training with the QRPO reference codebase (https://github.com/CLAIRE-Labo/quantile-reward-policy-optimization). Part of the dataset collection for the paper Quantile Reward Policy Optimization: Alignment with Pointwise Regression and Exact Partition Functions (https://arxiv.org/pdf/2507.08068). tabular10K<n<100K1 likes214 downloads10mo agoHugging Face12DenCT /leetcode-python-solutions-with-exaplanationstabulartext-generation10K<n<100K2 likes148 downloads2y agoHugging Face13ronantakizawa /leetcode-assembly LeetCode Assembly Dataset 441 LeetCode problems solved in C, compiled to assembly across 4 architectures, 2 compilers, and 4 optimization levels using GCC and Clang via the Godbolt Compiler Explorer API. Dataset Summary Stat Value Total rows 14,112 Unique problems 441 Architectures x86-64, AArch64, MIPS64, RISC-V 64 Compilers GCC 15.2, Clang 21.1.0 Optimization levels -O0, -O1, -O2, -O3 Compilation success rate 100% Difficulty split Easy: 98… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/leetcode-assembly.tabulartext-generation10K<n<100K12 likes91 downloads8mo agoHugging Face14minfeng-ai /leetcode_preference Dataset Card for LeetCode Preference Dataset Summary This dataset facilitates experiments utilizing Direct Preference Optimization (DPO) as outlined in the paper titled Direct Preference Optimization: Your Language Model is Secretly a Reward Model. This repository provides code pairings crafted by CodeLLaMA-7b. For every LeetCode question posed, CodeLLaMA-7b produces two unique solutions. These are subsequently evaluated and ranked by human experts based on their accuracy… See the full description on the dataset page: https://huggingface.co/datasets/minfeng-ai/leetcode_preference.tabularn<1K7 likes66 downloads3y agoHugging Face15ziwenyd /leetcode-standalone-wordstabular1K<n<10K0 likes60 downloads4y agoHugging Face16LimYeri /leetcode_with_youtube_captionsimagetext-classification10K<n<100K1 likes32 downloads2y agoHugging Face17BoyuanJackchen /leetcode_free_questions_labeledtabular1K<n<10K5 likes28 downloads4y agoHugging Face18NyanDoggo /leetcodeThis dataset contains python solutions for various Leetcode problems, scraped from different posts by users from the solutions tab. tabular1K<n<10K1 likes17 downloads2y agoHugging Face19nhyha /leetcode_with_youtube_captionsimage10K<n<100K0 likes16 downloads2y agoHugging Face20Inv /LeetCodeDataset-qwen3.5-0.8b-rl-bandtabularn<1K0 likes16 downloads2mo agoHugging Face21alexdzm /leetcode-no-depstabular1K<n<10K1 likes14 downloads1y agoHugging Face22mlfoundations-dev /b2_code_fasttext_pos_leetcode_neg_sqltabular10K<n<100K0 likes11 downloads1y agoHugging Face23mlfoundations-dev /b2_code_fasttext_pos_leetcode_neg_sql_3k_eval_636d mlfoundations-dev/b2_code_fasttext_pos_leetcode_neg_sql_3k_eval_636d Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces Accuracy 19.7 60.0 73.8 30.2 44.6 43.3 33.1 10.6 14.4 AIME24 Average Accuracy: 19.67% ± 2.42% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1 10.00% 3 30 2 33.33% 10 30 3… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/b2_code_fasttext_pos_leetcode_neg_sql_3k_eval_636d.tabular1K<n<10K0 likes11 downloads1y agoHugging Face24mlfoundations-dev /b2_code_fasttext_pos_leetcode_neg_sql_10ktabular10K<n<100K0 likes10 downloads1y agoHugging Face25wassname /vgrout-leetcode-teacher-demos vGROUT LeetCode teacher demonstrations Cached teacher demonstrations used to warm up the vGROUT gradient-routing experiments on the ariahw/rl-rewardhacking LeetCode environment. Each row is a full problem-specific completion. The kind column gives the two demonstration types: hack (215 rows): verified exploits of the run_tests loophole (hacked=True, gt_pass=False). solve (126 rows): correct solutions verified against the ground-truth tests (gt_pass=True). Why fewer… See the full description on the dataset page: https://huggingface.co/datasets/wassname/vgrout-leetcode-teacher-demos.tabulartext-generationn<1K0 likes10 downloads4mo agoHugging Face26mlfoundations-dev /b2_code_fasttext_pos_leetcode_neg_sql_0.3k_eval_636d mlfoundations-dev/b2_code_fasttext_pos_leetcode_neg_sql_0.3k_eval_636d Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces Accuracy 17.7 56.0 73.6 29.0 43.1 37.7 29.0 8.3 9.9 AIME24 Average Accuracy: 17.67% ± 1.77% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1 10.00% 3 30 2 16.67% 5 30 3… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/b2_code_fasttext_pos_leetcode_neg_sql_0.3k_eval_636d.tabular1K<n<10K0 likes9 downloads1y agoHugging Face27LimYeri /LeetCode_YT_CC_CoT_SummarygatedLeetCode Information & YouTube Captions with CoT Summaries Original data -> LimYeri/leetcode_with_youtube_captions The original ['cc_content'] column had tokens that were too long and contained a lot of repetition, which necessitated summarization. Consequently, our team (Project Team: CodeMind) summarized the ['cc_content'] column data using the Chain of Thought (CoT) technique with the gpt-3.5-turbo-0125 & gpt-4-turbo-2024-04-09 model. -> new column ['Summary'] imagetext-classification10K<n<100K1 likes7 downloads2y agoHugging Face28mlfoundations-dev /b2_code_fasttext_pos_leetcode_neg_sql_10k_eval_636d mlfoundations-dev/b2_code_fasttext_pos_leetcode_neg_sql_10k_eval_636d Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces Accuracy 18.7 59.5 77.8 28.6 42.9 39.9 38.5 11.7 16.2 AIME24 Average Accuracy: 18.67% ± 0.97% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1 20.00% 6 30 2 20.00% 6 30 3… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/b2_code_fasttext_pos_leetcode_neg_sql_10k_eval_636d.tabular1K<n<10K0 likes7 downloads1y agoHugging Face29mlfoundations-dev /b2_code_fasttext_pos_leetcode_neg_sql_1k_eval_636d mlfoundations-dev/b2_code_fasttext_pos_leetcode_neg_sql_1k_eval_636d Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces Accuracy 17.7 51.8 71.4 27.0 39.8 42.8 28.8 6.9 10.4 AIME24 Average Accuracy: 17.67% ± 1.16% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1 16.67% 5 30 2 23.33% 7 30 3… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/b2_code_fasttext_pos_leetcode_neg_sql_1k_eval_636d.tabular1K<n<10K0 likes7 downloads1y agoHugging Face30mlfoundations-dev /b2_code_fasttext_pos_leetcode_neg_sql_eval_636d mlfoundations-dev/b2_code_fasttext_pos_leetcode_neg_sql_eval_636d Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces Accuracy 19.3 62.7 76.2 26.6 41.5 41.9 43.7 14.6 17.2 AIME24 Average Accuracy: 19.33% ± 1.14% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1 16.67% 5 30 2 23.33% 7 30 3 20.00%… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/b2_code_fasttext_pos_leetcode_neg_sql_eval_636d.tabular1K<n<10K0 likes6 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.