Team Ai
Datasetpublic

sunweiwei/ai4sci-surface-code-decoding

Sycamore surface-code decoding: materialized benchmark Predict a logical observable flip from repeated stabilizer detection events in a noisy quantum memory. The benchmark trains decoders that improve the reliability of encoded quantum information. It uses real Sycamore hard-readout experiments at code distances 3 and 5, not simulated soft-readout d11 data. Source: Google Quantum AI Sycamore memory experiments, Zenodo 6804040, CC-BY-4.0. Scientific model reference: Bausch et… See the full description on the dataset page: https://huggingface.co/datasets/sunweiwei/ai4sci-surface-code-decoding.

sourceHugging Facecc-by-4.0updated 17d agoView on Hugging Face
0likes84downloads
README.md104 linesDownload Raw Back to root
1---2language:3- en4license: cc-by-4.05pretty_name: Sycamore Surface-code Decoding — AI4Sci Materialized Release6size_categories:7- 1M<n<10M8task_categories:9- tabular-classification10tags:11- quantum-error-correction12- surface-code13- sycamore14- alphaqubit15- ai4science16---17 18# Sycamore surface-code decoding: materialized benchmark19 20Predict a logical observable flip from repeated stabilizer detection events in21a noisy quantum memory. The benchmark trains decoders that improve the22reliability of encoded quantum information. It uses real Sycamore hard-readout23experiments at code distances 3 and 5, not simulated soft-readout d11 data.24 25![Scientific problem, real validation example and output](science-introduction.png)26 27Source: [Google Quantum AI Sycamore memory experiments, Zenodo 6804040](https://zenodo.org/records/6804040),28CC-BY-4.0. Scientific model reference:29[Bausch et al., Learning high-accuracy error decoding for quantum processors,30Nature 635, 834–840 (2024)](https://www.nature.com/articles/s41586-024-08148-8).31The reference AlphaQubit implementation is an independent reproduction, not32official Google code or weights, and does not reproduce the paper's headline33ensemble accuracy. Its trained weights are not distributed to agents.34 35## Contents and splits36 37- `development.tar.gz`: physically separate train and validation arrays,38  ideal circuits, published even-fitted detector error models, and `reference/`:39  the reference model's validation predictions, validation metrics, aggregate test40  score and training cost, plus a submission template without weights.41- `verifier-inputs.tar.gz`: input-only test arrays and ideal circuits. No logical42  labels, raw measurement records, historical decoder predictions, or odd-fit43  noise models are in this archive.44- `release-manifest.json`: archive sizes, expanded sizes and SHA-256 identities.45- `NOTICE.md`: attribution and source/licensing notes.46 47Each of 130 conditions contains 50,000 original shots: four d3 patches and one48d5 patch, X/Z bases, and 13 odd round counts 1..25. Original zero-based rows are49split into first 19,880 even rows for train, last 5,120 even rows for validation,50and all 25,000 odd rows for test. Totals are 2,584,400 / 665,600 / 3,250,000.51There is no geometry or acquisition-session holdout. IDs are opaque uint6452values; the inference interface does not expose original source row numbers.53 54Training/validation NPZs contain `sample_ids:uint64[N]`, packed55`detectors:uint8[N,ceil(D/8)]` (little-endian bit order), and `labels:uint8[N]`.56Verifier inputs contain only the first two. A split manifest specifies geometry,57array paths, ideal circuits, shot counts and hashes. Targets are the official58logical-path observable flips, not raw final-data-qubit parity. Noise DEMs are59calibrated using all even syndromes, including validation syndromes, but no60logical labels or odd test data. This is a declared calibration exception.61 62The original dataset and historical test are public. A separate private63`Corning/ai4sci-surface-code-decoding-evaluation` repository supplies labels to64trusted benchmark operators; this runtime separation does not make the source65historically secret. Never expose test roles to an agent's development runtime.66 67## Model and score68 69The reference model has a recurrent Transformer core with width 320, three70blocks per round, four attention heads, and geometry/event-dependent bias.71Eight d3 specialists have 8,449,990 parameters each; two d5 specialists have728,456,070 each; total 84,512,060. Exactly one specialist serves each shot.73Weights are plain FP32 NumPy arrays, with BF16 compute in the frozen configuration.74 75Overall P = 1 minus balanced logical failure: average durations in each76geometry/basis group, groups within distance, and give d3/d5 equal weight.77Normalized score = clip((P-L)/(U-L),0,1), with fixed L=0.7160463461538461 from78MWPM and U=0.7615159615384615 from the strongest same-test released79tensor-network predictions. TN is a rescored historical reference, not a80newly reimplemented decoder. The reference AlphaQubit gives P=0.755060769230769381and normalized score=0.8580328368056459. Raw/unclipped scores and scientific LER,82error suppression and calibration metrics remain available.83 84The original accepted neural training/replay cost was about 200.3 allocated85GPU-hours on RTX PRO 6000 Blackwell Max-Q 96 GB hardware; this is not an H10086throughput claim. Agents train from scratch; the agent budget is under review.87 88## Reproduce89 90See the [GitHub task](https://github.com/T0-RSI/ai4sci-tasks/tree/main/surface-code-decoding)91for `environment/data/materialize.py`, the immutable HF lock, Harbor environments,92strict submission schema and trusted verifier. Normal deployment downloads93these pre-materialized archives; it does not regenerate experimental data.94Use the full commit revision and archive hashes pinned in GitHub. Agent data,95inference inputs and trusted labels must be mounted separately. The verifier96runs without network, replays a relocated trained model, seals outputs, and97scores them in a separate label-holding service.98 99Do not compare the normalized scalar as a universal measure of scientific100difficulty across tasks. It measures progress between declared task-specific101references and saturates beyond the upper reference; raw metrics preserve102further progress. The scientific evaluation is limited to this historical103hard-readout d3/d5 distribution.104