neulab/leetcode
LeetCode multilingual benchmark dataset LeetCode problems for the msl-multilingual-self-learning benchmark, flattened to one row per (problem, language) in 9 languages, stored as data/<split>/<lang>-NNN.jsonl. Each row has the interface for its language (from LeetCode's code snippets), the shared canonical_tests (Python asserts from newfacade/LeetCodeDataset), the problem_description (from the LeetCode page) and metadata. Tests that break the problem's Constraints, do not fit a… See the full description on the dataset page: https://huggingface.co/datasets/neulab/leetcode.
LeetCode multilingual benchmark dataset
LeetCode problems for the msl-multilingual-self-learning benchmark, flattened to one row per (problem, language) in 9 languages, stored as data/<split>/<lang>-NNN.jsonl. Each row has the interface for its language (from LeetCode's code snippets), the shared canonical_tests (Python asserts from newfacade/LeetCodeDataset), the problem_description (from the LeetCode page) and metadata.
Tests that break the problem's Constraints, do not fit a declared type in some language, or expect inf/nan were removed; problems that take trees or linked lists, modify their input in place, or kept fewer than 10 tests were dropped. The test split also drops problems with several valid answers or whose text refers to a figure. Where answers may come in any order or are decimals, the asserts call answers_match (defined at the top of the test source), and canonical_tests.comparison says which rule applies ("unordered", "float" or "exact"). public_tests holds the inputs of LeetCode's Testcase panel (one string per case, one JSON value per line in parameter order; LeetCode publishes no outputs), and hints the page's hints as plain text. reports/dropped.md lists every dropped problem and test count, and reports/<split>.json has the details.
Splits
train: 2103 problems x 9 languages = 18927 rows (2641 in the source dataset)test: 191 problems x 9 languages = 1719 rows (228 in the source dataset)
Load
from datasets import load_dataset
train = load_dataset("neulab/leetcode", split="train")
test = load_dataset("neulab/leetcode", split="test")
python_rows = test.filter(lambda r: r["language"] == "python")