Team Ai
Datasetpublic

smshahbaj/verifiable-code-reasoning

Verifiable Code Reasoning Execution-verified Python problems with chain-of-thought Sandbox-checked solutions · Multi-test unit checks · Deduplicated instances · Training-ready sft_text Overview Verifiable Code Reasoning is a large-scale dataset of Python coding problems where every kept solution has passed sandboxed unit tests. Unlike scraped contest dumps or unverified LLM traces, an example enters this release only if: a reference… See the full description on the dataset page: https://huggingface.co/datasets/smshahbaj/verifiable-code-reasoning.

sourceHugging Facemitupdated 20d agoView on Hugging Face
2likes1.8kdownloads
Dataset Card

<div align="center">

Verifiable Code Reasoning

Execution-verified Python problems with chain-of-thought

Sandbox-checked solutions · Multi-test unit checks · Deduplicated instances · Training-ready `sft_text`

![License: MIT](https://opensource.org/licenses/MIT) ![Examples](https://huggingface.co/datasets/smshahbaj/verifiable-code-reasoning)

</div>


Overview

Verifiable Code Reasoning is a large-scale dataset of Python coding problems where every kept solution has passed sandboxed unit tests.

Unlike scraped contest dumps or unverified LLM traces, an example enters this release only if:

  1. 1.a reference implementation exists,
  2. 2.it runs under a time-limited sandbox, and
  3. 3.it matches expected outputs on at least 3 tests.

Failed executions are discarded. Near-duplicate instances are filtered by content fingerprint (signature + tests), so shared solution templates across different inputs are allowed, but repeated test instances are not.

PropertyDetail
VerificationMultiprocess sandbox, timeout, ≥3 tests
DedupUnique id + instance content hash (signature + tests)
BalanceCategory & difficulty quotas during generation
ResumeHub checkpoints: train.jsonl, progress.json
LicenseMIT

Current release size: 1,500,000 verified examples (generation target 1,500,000).


Dataset statistics

Difficulty

DifficultyCount
easy783,088
medium698,789
hard18,123

Category

CategoryCount
array313,191
hashmap240,984
two_pointers218,129
greedy144,377
binary_search134,400
dp134,400
string133,673
bit132,948
math46,941
stack702
matrix255

Schema

FieldTypeDescription
idstringStable example id
problemstringNatural-language problem statement
signaturestringPython function signature
reasoningstringNumbered chain-of-thought
codestringReference Python solution
testsstringJSON list of {"input": ..., "output": ...} (≥3)
categorystringarray, hashmap, dp, binary_search, ...
difficultystringeasy / medium / hard
sft_textstringReady-to-train prompt block
verifiedboolAlways true in this release
code_hashstringHash of solution source (audit)
languagestringpython
tests is stored as a JSON string so Arrow/HF can keep one schema while outputs may be int, bool, or list.

Quick start

python
from datasets import load_dataset
import json

ds = load_dataset("smshahbaj/verifiable-code-reasoning")
row = ds["train"][0]
print(row["sft_text"][:500])
tests = json.loads(row["tests"])
print(tests[0])

SFT format (sft_text)

text
### Problem
...

### Reasoning
1) ...
2) ...

### Solution

...


How it was built

  1. 1.25 algorithmic generator families (arrays, strings, math, DP, greedy, binary search, bits, matrices, …)
  2. 2.Each example gets ≥3 unit tests (including varied inputs)
  3. 3.Solution is executed in a sandbox; mismatch / timeout / exception → drop
  4. 4.Quotas limit over-representation of a single category or difficulty
  5. 5.Instance dedup on signature+tests fingerprint (unrepeatable test cases)
  6. 6.Multi-session scale via Hugging Face checkpoints every 2000 new verifies

Intended uses

  • —Supervised fine-tuning on code + reasoning
  • —Process / outcome supervision signals (execution as ground truth)
  • —Filtering or ranking other synthetic coding traces

Out of scope / limitations

  • —Coverage is bounded by the generator bank (not full competitive-programming breadth)
  • —Python only in this release
  • —Synthetic style can still be distributionally narrow — always evaluate models on real benchmarks (HumanEval, MBPP, LiveCodeBench, etc.)
  • —Not a substitute for human-written contest editorials

Citation

bibtex
@misc{verifiable_code_reasoning,
  title        = {Verifiable Code Reasoning},
  author       = {smshahbaj},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/datasets/smshahbaj/verifiable-code-reasoning}},
  note         = {Execution-verified Python coding CoT dataset}
}

<div align="center">

Verified or it does not ship.

</div>