smshahbaj/verifiable-code-reasoning
Verifiable Code Reasoning Execution-verified Python problems with chain-of-thought Sandbox-checked solutions · Multi-test unit checks · Deduplicated instances · Training-ready sft_text Overview Verifiable Code Reasoning is a large-scale dataset of Python coding problems where every kept solution has passed sandboxed unit tests. Unlike scraped contest dumps or unverified LLM traces, an example enters this release only if: a reference… See the full description on the dataset page: https://huggingface.co/datasets/smshahbaj/verifiable-code-reasoning.
<div align="center">
Verifiable Code Reasoning
Execution-verified Python problems with chain-of-thought
Sandbox-checked solutions · Multi-test unit checks · Deduplicated instances · Training-ready `sft_text`
 
</div>
Overview
Verifiable Code Reasoning is a large-scale dataset of Python coding problems where every kept solution has passed sandboxed unit tests.
Unlike scraped contest dumps or unverified LLM traces, an example enters this release only if:
- a reference implementation exists,
- it runs under a time-limited sandbox, and
- it matches expected outputs on at least 3 tests.
Failed executions are discarded. Near-duplicate instances are filtered by content fingerprint (signature + tests), so shared solution templates across different inputs are allowed, but repeated test instances are not.
Current release size: 1,500,000 verified examples (generation target 1,500,000).
Dataset statistics
Difficulty
Category
Schema
tests is stored as a JSON string so Arrow/HF can keep one schema while outputs may be int, bool, or list.Quick start
from datasets import load_dataset
import json
ds = load_dataset("smshahbaj/verifiable-code-reasoning")
row = ds["train"][0]
print(row["sft_text"][:500])
tests = json.loads(row["tests"])
print(tests[0])SFT format (sft_text)
### Problem
...
### Reasoning
1) ...
2) ...
### Solution...
How it was built
- 25 algorithmic generator families (arrays, strings, math, DP, greedy, binary search, bits, matrices, …)
- Each example gets ≥3 unit tests (including varied inputs)
- Solution is executed in a sandbox; mismatch / timeout / exception → drop
- Quotas limit over-representation of a single category or difficulty
- Instance dedup on signature+tests fingerprint (unrepeatable test cases)
- Multi-session scale via Hugging Face checkpoints every 2000 new verifies
Intended uses
- Supervised fine-tuning on code + reasoning
- Process / outcome supervision signals (execution as ground truth)
- Filtering or ranking other synthetic coding traces
Out of scope / limitations
- Coverage is bounded by the generator bank (not full competitive-programming breadth)
- Python only in this release
- Synthetic style can still be distributionally narrow — always evaluate models on real benchmarks (HumanEval, MBPP, LiveCodeBench, etc.)
- Not a substitute for human-written contest editorials
Citation
@misc{verifiable_code_reasoning,
title = {Verifiable Code Reasoning},
author = {smshahbaj},
year = {2026},
howpublished = {\url{https://huggingface.co/datasets/smshahbaj/verifiable-code-reasoning}},
note = {Execution-verified Python coding CoT dataset}
}<div align="center">
Verified or it does not ship.
</div>
