Team Ai
Datasetpublic

ronantakizawa/leetcode-assembly

LeetCode Assembly Dataset 441 LeetCode problems solved in C, compiled to assembly across 4 architectures, 2 compilers, and 4 optimization levels using GCC and Clang via the Godbolt Compiler Explorer API. Dataset Summary Stat Value Total rows 14,112 Unique problems 441 Architectures x86-64, AArch64, MIPS64, RISC-V 64 Compilers GCC 15.2, Clang 21.1.0 Optimization levels -O0, -O1, -O2, -O3 Compilation success rate 100% Difficulty split Easy: 98… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/leetcode-assembly.

sourceHugging Faceapache-2.0updated 8mo agoView on Hugging Face
12likes90downloads
Dataset Card

LeetCode Assembly Dataset

441 LeetCode problems solved in C, compiled to assembly across 4 architectures, 2 compilers, and 4 optimization levels using GCC and Clang via the Godbolt Compiler Explorer API.

Dataset Summary

StatValue
Total rows14,112
Unique problems441
Architecturesx86-64, AArch64, MIPS64, RISC-V 64
CompilersGCC 15.2, Clang 21.1.0
Optimization levels-O0, -O1, -O2, -O3
Compilation success rate100%
Difficulty splitEasy: 98, Medium: 259, Hard: 84

Each problem has 32 assembly variants (4 architectures x 2 compilers x 4 optimization levels), making this useful for studying how the same algorithm compiles differently across ISAs, compilers, and optimization settings.

Key Insights

  • —-O3 often produces more code than -O0, not less. Across all architectures, -O3 output is 37-73% larger on average due to aggressive inlining and loop unrolling.
  • —-O1 consistently produces the most compact assembly across all four architectures.
  • —Recursive tree problems see extreme -O3 blowup. The top 10 largest O3-vs-O0 ratios are all binary tree problems. Worst case: Longest Univalue Path goes from 90 lines (O0) to 3,809 lines (O3) -- a 42x increase.
  • —Iterative and math problems shrink dramatically at -O3. Flip String to Monotone Increasing drops from 89 to 8 lines (91% reduction). Nim Game compiles down to just 4 lines.
  • —AArch64 produces the most compact assembly at -O2 (avg 153 lines), followed by x86-64 (170), RISC-V (170), and MIPS64 (236). MIPS is 39% more verbose than AArch64.
  • —Hard problems produce 2x more assembly than Easy problems at -O2 on x86-64 (247 vs 121 avg lines). Algorithmic complexity maps directly to code size.
  • —The largest output is LFU Cache (1,121 lines at -O0 on x86-64), a complex doubly-linked list plus hash table implementation.
  • —The smallest output is Nim Game at -O3 (4 lines on x86-64) -- the entire solution optimizes to a single bitwise AND.

Use Cases

  • —Training or evaluating models on C-to-assembly translation
  • —Studying how optimization levels affect generated code
  • —Compiler behavior research

Schema

ColumnTypeDescription
idintUnique row identifier (0-14111)
problem_idintLeetCode problem number
problem_titlestringProblem name (e.g. "Two Sum")
difficultystringEasy, Medium, or Hard
c_sourcestringComplete C source code
architecturestringTarget ISA: x86-64, aarch64, mips64, riscv64
optimizationstringOptimization flag: -O0, -O1, -O2, -O3
compilerstringCompiler version used (e.g. x86-64 gcc 15.2, x86-64 clang 21.1.0)
assemblystringCompiled assembly output

Usage

python
from datasets import load_dataset

ds = load_dataset("ronantakizawa/leetcode-assembly", split="train")

# Get all x86-64 GCC assembly at -O2
x86_gcc_O2 = ds.filter(lambda r: r["architecture"] == "x86-64" and r["compiler"].startswith("x86-64 gcc") and r["optimization"] == "-O2")

# Compare GCC vs Clang for the same problem
two_sum = ds.filter(lambda r: r["problem_id"] == 1 and r["architecture"] == "x86-64" and r["optimization"] == "-O2")
for row in two_sum:
    print(f"--- {row['compiler']} ---")
    print(row["assembly"][:200])
    print()

Sources

C solutions were collected from three open-source GitHub repositories of pure-C LeetCode solutions:

After deduplication, 441 unique problems remained. Solutions are self-contained C files using only the standard library (no C++ STL), producing clean and readable assembly output.