Team Ai
Datasetpublic

GPUMODE/one-layer-deeper-submissions

One Layer Deeper submissions This dataset archives 15,602 accepted uploads from 206 GitHub accounts to the One Layer Deeper competition. It contains 9,627 distinct source files, all upload metadata, and stored evaluation results. Snapshot: September 7, 2026, 21:48 UTC, after the August 31 submission deadline. Split Uploads Succeeded Failed easy 11,961 11,112 849 medium 2,704 2,509 195 hard 937 847 90 All accepted uploads are included: practice runs, failures… See the full description on the dataset page: https://huggingface.co/datasets/GPUMODE/one-layer-deeper-submissions.

sourceHugging Faceupdated 1mo agoView on Hugging Face
1likes174downloads
Dataset Card

One Layer Deeper submissions

This dataset archives 15,602 accepted uploads from 206 GitHub accounts to the One Layer Deeper competition. It contains 9,627 distinct source files, all upload metadata, and stored evaluation results. Snapshot: September 7, 2026, 21:48 UTC, after the August 31 submission deadline.

SplitUploadsSucceededFailed
easy11,96111,112849
medium2,7042,509195
hard93784790

All accepted uploads are included: practice runs, failures, identical resubmissions, and superseded or excluded entries. “Succeeded” is the evaluator's run status, not a claim of task correctness or rule compliance. GitHub accounts do not necessarily represent distinct people. Rejected requests and local experiments are outside this archive.

Load the data

This is a private dataset under GPUMODE. Access requires an authorized Hugging Face account. Run hf auth login first; datasets uses that local authentication for the download.

python
from datasets import load_dataset

hard = load_dataset("GPUMODE/one-layer-deeper-submissions", split="hard")
row = hard[0]
print(row["id"], row["github_login"], row["status"])
print(row["source"][:500])

Install the datasets package to use this example. Loading reads participant source as text; executing that source is unnecessary for archive analysis.

Each row has these fields:

FieldMeaning
idAccepted submission UUID
tiereasy, medium, or hard
github_loginSubmitting account label
created_atStored upload timestamp, including UTC offset
statusStored evaluator run status
sha256SHA-256 of the original UTF-8 source bytes
source_bytesOriginal source length in bytes
leaderboard_rankRank of this exact upload in the archived public leaderboard, or null
sourceComplete original submission.py decoded as UTF-8, preserving line endings
metadata_jsonComplete original metadata.json file as a UTF-8 string, preserving formatting
result_jsonComplete original result.json file as a UTF-8 string, preserving formatting

A failed run may have JSON null in result_json; use json.loads(row["result_json"]) to interpret it. Scores are recorded evaluator outcomes, not independent reproductions. A high score does not establish rule compliance or generalization beyond the recorded tests.

Original files and integrity

text
data/{easy,medium,hard}.parquet   One row per accepted upload
archives/{easy,medium,hard}.tar.gz
  submissions/<tier>/<id>/submission.py
  submissions/<tier>/<id>/metadata.json
  submissions/<tier>/<id>/result.json
inventory/manifest.json
inventory/export_summary.json
inventory/statistics.json
inventory/engagement.json
inventory/leaderboard.json
verification.json               Local packaging verification report
checksums.json                  File SHA-256 hashes and sizes
SHA256SUMS                      Checksums of every other packaged file

The tar archives preserve all three files byte for byte. Archive order and timestamps are normalized; the files' content bytes are unchanged. Parquet strings also round-trip exactly to the original bytes via UTF-8 encoding, including CRLF line endings and JSON whitespace.

python
import hashlib
import json

source_bytes = row["source"].encode("utf-8")
assert len(source_bytes) == row["source_bytes"]
assert hashlib.sha256(source_bytes).hexdigest() == row["sha256"]
metadata = json.loads(row["metadata_json"])
result = json.loads(row["result_json"])

Packaging verified every Parquet source hash, every metadata/result string against its original bytes, every archive entry against its original file, and exact submission-ID coverage in each split. checksums.json maps every substantive file path to its SHA-256 and byte length, excluding the two checksum manifests. SHA256SUMS additionally covers checksums.json; it does not hash itself. With the repository downloaded locally, run sha256sum -c SHA256SUMS (or shasum -a 256 -c SHA256SUMS on macOS).

Provenance and scope

The original export used a read-only, repeatable-read transaction against the organizer database. Every source matched the database's md5(source) value, and SHA-256 was saved per upload. The public leaderboard was fetched separately immediately before that transaction. inventory/export_summary.json records the snapshot and exclusions; inventory/manifest.json retains per-upload provenance.

This release explicitly includes only the submission files, five listed inventory files, and packaging documentation. It does not include dataset inputs, trained checkpoints, raw service logs, metric histories, moderation records, credentials, or email fields. Participant source is retained as submitted. Moderation status is not a row-level classification in this release, and absence from the leaderboard alone does not establish a reason for exclusion.

Participant files retain their existing ownership and terms. No blanket license or relicensing is asserted for these uploads. The upstream service's Apache license does not automatically apply to participant submissions.

The packaging script reads code as bytes and text and never imports or executes participant source. Rebuilding with the same inputs and Python/PyArrow versions produces deterministic archive content and packaging metadata; the exact runtime versions are recorded in verification.json.