datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LeetCodeDataset
LeetCodeDataset
LeetCodeDataset is a dataset consists of Python leetcode problems that can be used for LLM training and evaluation.
💻 GitHub
📄 LeetCodeDataset: A Temporal Dataset for Robust Evaluation and Efficient Training of Code LLMs
📄 Policy Filtration for RLHF to Mitigate Noise in Reward Models
leetcode-problem-solutions
LeetCode Solution Dataset
This dataset contains community-contributed LeetCode solutions scraped from public discussions and solution pages, enriched with metadata such as vote counts, author info, tags, and full code content. The goal is to make high-quality, peer-reviewed coding solutions programmatically accessible for research, analysis, educational use, or developer tooling.
Column Descriptions
Column Name
Type
Description
question_slug
string
The unique… See the full description on the dataset page: https://huggingface.co/datasets/kaysss/leetcode-problem-solutions.leetcodetw-leetcode
Dataset Card for tw-leetcode
A curated Traditional Chinese LeetCode solution dataset with high-efficiency answers (Beats 100%), structured explanation in "Top Concept → Step Implement → Complexity Analysis" style, updated daily.
Dataset Details
Dataset Description
tw-leetcode 是一個針對 LeetCode 題目的繁體中文資料集,內容包含高效能程式解法、完整的解題思路,以及時間與空間複雜度分析。每份題解都經由人工清洗與優化,並依循「Top Concept → Step Implement → Complexity Explanation」的結構撰寫,方便機器學習模型或人類讀者理解程式邏輯的推理過程。
本資料集適合作為:… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/tw-leetcode.deepseek-leetcodeDeepseek Leetcode dataset from https://github.com/deepseek-ai/DeepSeek-Coder/tree/main/Evaluation/LeetCode
leetcode-problem-set
LeetCode Scraper Dataset
This dataset contains information scraped from LeetCode. It is designed to assist developers in analyzing LeetCode problems, generating insights, and building tools for competitive programming or educational purposes.
Dataset Contents
The dataset includes the following files:
problem_set.csv
Contains a list of LeetCode problems with metadata such as difficulty, acceptance rate, tags, and more.
Columns:
acRate: Acceptance rate of the… See the full description on the dataset page: https://huggingface.co/datasets/kaysss/leetcode-problem-set.leetcode_pythondoocs-leetcode-solutions
Doocs LeetCode Solutions
LeetCode problems with solutions in many programming languages, built from the Doocs LeetCode repository. Each solution comes with the approach name, the reasoning that leads to it, and an explanation with complexity analysis. The dataset is meant for fine-tuning and evaluating code generation models.
The dataset is regenerated monthly from the latest Doocs commit by the leetcode-dataset-generator tool.
Structure
The dataset has two… See the full description on the dataset page: https://huggingface.co/datasets/olegshulyakov/doocs-leetcode-solutions.leetcode-problem-detailed
LeetCode Scraper Dataset
This dataset contains information scraped from LeetCode, including problem details, metadata, and related files. It is designed to assist developers in analyzing LeetCode problems, generating insights, and building tools for competitive programming or educational purposes.
questions_deets.csv
Contains detailed information about each problem, including problem descriptions, constraints, and examples.
Columns:
questionFrontendId: Unique problem ID.… See the full description on the dataset page: https://huggingface.co/datasets/kaysss/leetcode-problem-detailed.grpo-qwen3-1.7b-nemotron-leetcode-clean-3.2k-bs32-n8-verl091-epoch2-146102-rollouts
Coding RL rollouts
grpo_Qwen3-1.7B_Nemotron-LeetCode-clean-3.2k_bs32_n8_seqs16_32k_epoch2_verl091
One verified gzip JSONL shard per training step; 256 responses per shard.
LCB binary grading after thinking, without an EOS gate.
leetcode
LeetCode multilingual benchmark dataset
LeetCode problems for the msl-multilingual-self-learning benchmark, flattened to
one row per (problem, language) in 9 languages, stored as data/<split>/<lang>-NNN.jsonl.
Each row has the interface for its language (from LeetCode's code snippets),
the shared canonical_tests (Python asserts from newfacade/LeetCodeDataset),
the problem_description (from the LeetCode page) and metadata.
Tests that break the problem's Constraints, do not fit a… See the full description on the dataset page: https://huggingface.co/datasets/neulab/leetcode.Nemotron-LeetCode-coding-clean-3.2k
Nemotron + LeetCode Coding Clean 3.2k
3,200 distinct training problems, seed 42, intended for Python coding reinforcement learning. This is a training mix, not a held-out benchmark. It combines the pinned default train Parquet split of Nemotron-RL-coding-competitive_coding with the train JSONL of LeetCodeDataset.
Composition
Source
Questions
Selection
Nemotron / Codeforces
1,856
1000–1600, inclusive
Nemotron / AtCoder
366
300–2000 display difficulty… See the full description on the dataset page: https://huggingface.co/datasets/hi-todayis-jh/Nemotron-LeetCode-coding-clean-3.2k.tigerbot-kaggle-leetcodesolutions-en-2kTigerbot 基于leetcode-solutions数据集,加工生成的代码类sft数据集
原始来源:https://www.kaggle.com/datasets/erichartford/leetcode-solutions
Usage
import datasets
ds_sft = datasets.load_dataset('TigerResearch/tigerbot-kaggle-leetcodesolutions-en-2k')
grpo-qwen3-1.7b-nemotron-leetcode-clean-3.2k-bs32-n8-verl091-146102-rollouts
Coding GRPO rollouts
grpo_Qwen3-1.7B_Nemotron-LeetCode-clean-3.2k_bs32_n8_seqs16_32k_1epoch_verl091
One verified gzip JSONL shard per training step; 256 responses per shard.
LCB binary grading after thinking, without an EOS gate.
leetcode-complete
Complete LeetCode Problems Dataset
This dataset contains a comprehensive collection of LeetCode problems (including premium) with AI-generated solutions in JSONL format. It is regularly updated to include new problems as they are added to LeetCode.
Splits
The dataset is divided into the following splits:
train: Contains approximately 80% of the problems for training
validation: Contains approximately 10% of the problems for validation
test: Contains approximately… See the full description on the dataset page: https://huggingface.co/datasets/whiskwhite/leetcode-complete.qrpo-paper-llama-nosft-leetcode-sandbox-temp1-ref50-offpolicy10random-sandbox
qrpo-paper-llama-nosft-leetcode-sandbox-temp1-ref50-offpolicy10random-sandbox
Dataset with reference completions and rewards for a specific model and reward model, ready for training with the QRPO reference codebase (https://github.com/CLAIRE-Labo/quantile-reward-policy-optimization).
Part of the dataset collection for the paper Quantile Reward Policy Optimization: Alignment with Pointwise Regression and Exact Partition Functions (https://arxiv.org/pdf/2507.08068).
leetcode1000leetcodeleetcode-standaloneleetcode_code_generationleetcode-solutionsFrom: https://www.kaggle.com/datasets/jacobhds/leetcode-solutions-and-content-kpis
LeetCode-OThis is the LeetCode-O benchmark proposed in CodeI/O paper (Arxiv 2502.07316).
The data file is in leetcode.jsonl and we provide an example prediction file (gpt-4.1-nano) in prediction.jsonl.
To evaluate your model on this benchmark, please prepare your outputs as in the format of prediction.jsonl, which is to add an output field to each line in leetcode.jsonl corresponding to the output of a LLM to the messages.
To calculate the scores, please simply follow evaluate.py, and you will see the… See the full description on the dataset page: https://huggingface.co/datasets/hkust-nlp/LeetCode-O.LeetCode-Contest
LeetCode Contest Benchmark
A new benchmark for evaluating Code LLMs proposed by DeepSeek-Coder, which consists of the latest algorithm problems of different difficulties.
Usage
git clone https://github.com/deepseek-ai/DeepSeek-Coder.git
cd Evaluation/LeetCode
# Set the model or path here
MODEL="deepseek-ai/deepseek-coder-7b-instruct"
python vllm_inference.py --model_name_or_path $MODEL --saved_path output/20240121-Jul.deepseek-coder-7b-instruct.jsonl
python… See the full description on the dataset page: https://huggingface.co/datasets/TechxGenus/LeetCode-Contest.qrpo-paper-llama-sft-leetcode-sandbox-temp1-ref50-offpolicy10random-sandbox
qrpo-paper-llama-sft-leetcode-sandbox-temp1-ref50-offpolicy10random-sandbox
Dataset with reference completions and rewards for a specific model and reward model, ready for training with the QRPO reference codebase (https://github.com/CLAIRE-Labo/quantile-reward-policy-optimization).
Part of the dataset collection for the paper Quantile Reward Policy Optimization: Alignment with Pointwise Regression and Exact Partition Functions (https://arxiv.org/pdf/2507.08068).
LeetCodeMetaDataleetcodesolutions_en_2kleetcode-python-dataset
leetcode-python-dataset
Code for building and publishing the justindal/leetcode-python-dataset dataset on Hugging Face.
Merges two open-source LeetCode datasets into a unified schema with consistent formatting, field normalisation, and solution validation.
Dataset
Split
Rows
Source
train
2856
newfacade + greengerong
valid
310
slug-group split from train
test
228
newfacade only
Schema
default config (training)
Each row is a… See the full description on the dataset page: https://huggingface.co/datasets/justindal/leetcode-python-dataset.leetcode_problem_solutionThis dataset contains: problems and solutions in Leetcode, crawled from: https://github.com/AnasImloul/Leetcode-Solutions
The format of data:
title: title of the problem
algo_input: the description of the problem
solution_py: the solution in Python
solution_js: the solution in Js
solution_java: the solution in Java
solution_c: the solution in C
leetcode
License & Attribution
MTEB-format derivative of greengerong/leetcode. Query = problem title + statement; document = Python solution. Licensed under MIT (same as source).
leetcode-python-solutions-with-exaplanations
