Terrano09/generalized-cpp-kernel-GRPO20
0
1---2base_model: zai-org/GLM-4.7-Flash3library_name: peft4tags:5- lora6- grpo7- code8- cpp9---10 11# Generalized C++ kernel GRPO2012 13Portable PEFT LoRA checkpoints from `generalized-cpp-kernel-grpo20-spot-20260829-083214-retry1`.14 15## Selected checkpoint16 17The selected release checkpoint is `iter_0000014`. Selection basis: higher four-trial Fixed26 Pass@1 and turn-2 means than iter 19.18 19| Field | Value |20| --- | --- |21| Base model | `zai-org/GLM-4.7-Flash@7dd20894a642a0aa287e9827cb1a1f7f91386b67` |22| Adapter SHA-256 | `b4bb3a250e28696c597d84db459caa75978e160996818dbfce22b8896b2c794b` |23| LoRA rank / alpha | 16 / 32 |24| Target modules | `q_a_proj`, `kv_a_proj_with_mqa`, `o_proj`, `gate_proj`, `up_proj`, `down_proj` |25| Planned updates | 20 |26| Run outcome | Retry completed all 20 updates and preserved checkpoints through iter 19. |27| Training data | `Generalized_CPP_GRPO20_train.jsonl`, 6 rows |28| Training-data SHA-256 | `566ea43aad4f6419b345e0295288c27b374575ce4db30cbd2d6b7510bf86ffbe` |29| Training-manifest SHA-256 | `25d6165985c440e0bc79f3a3f2a6aeeac4943324f2289bc03fd1d775d30bf3a6` |30 31## Post-training evaluations32 33Each row uses only its selected best four receipt-verified trials (26 tasks per trial, 104 task evaluations). Iterations are reported separately.34 35| Checkpoint | Pass@1 trial scores | Pass@1 mean | Turn-2 trial scores | Turn-2 mean |36| --- | --- | ---: | --- | ---: |37| `iter_0000014` | 13, 13, 11, 10 | 11.75/26 | 17, 17, 14, 16 | 16/26 |38| `iter_0000019` | 12, 10, 10, 9 | 10.25/26 | 17, 13, 13, 11 | 13.5/26 |39 40| Checkpoint | Pass@1 SD; range; task-bootstrap 95% CI (out of 26) | Turn-2 SD; range; task-bootstrap 95% CI (out of 26) | Conditional turn-2 recovery |41| --- | --- | --- | --- |42| `iter_0000014` | 1.50; 10-13; 8-15.5 | 1.41; 14-17; 11.75-20 | 17/57 (29.8%; CI 14.5-48.9%) |43| `iter_0000019` | 1.26; 9-12; 6.25-14.5 | 2.52; 11-17; 9.25-17.75 | 13/63 (20.6%; CI 9.5-35.0%) |44 45Evaluation used `fixed26-contract-v2`, thinking enabled, temperature 0.7, top-p 1.0, and a 32,768-token response limit. The complete selected run IDs and byte-for-byte receipts are under `evaluations/`.46 47## Training data48 496 executable shadow-task rows: clock, complex-numbers, crypto-square, grade-school, kindergarten-garden, and perfect-numbers. The manifest explicitly records overlap with six Fixed26 task IDs; reference answers are not packaged.50 51The exact JSONL and its source manifest are included at the repository root. Their hashes are checked during release construction.52 53## Checkpoints54 55Every checkpoint directory contains only the two portable inference artifacts: `adapter_config.json` and `adapter_model.bin`. Megatron tensor-parallel shards, optimizer state, and other training-only files are intentionally omitted.56 57| Checkpoint | Adapter SHA-256 |58| --- | --- |59| `iter_0000004` | `cf25cdb386b4e9cbc372480768882dc9f828b5b68e9017868c0255945d30aec6` |60| `iter_0000009` | `83eec62a120142917902355a641043cf1970658170056b6fce5ebf98dd1de7dd` |61| `iter_0000014` | `b4bb3a250e28696c597d84db459caa75978e160996818dbfce22b8896b2c794b` |62| `iter_0000019` | `e915a82a09a290cd4a6b587fe08fdad8acbbf58dea27e1a53e9e77f9effdf300` |63 64## Loading65 66```python67from peft import PeftModel68from transformers import AutoModelForCausalLM, AutoTokenizer69 70base = "zai-org/GLM-4.7-Flash"71checkpoint = "Terrano09/generalized-cpp-kernel-GRPO20"72subfolder = "checkpoints/iter_0000014/adapter"73 74tokenizer = AutoTokenizer.from_pretrained(base, trust_remote_code=True)75model = AutoModelForCausalLM.from_pretrained(base, trust_remote_code=True)76model = PeftModel.from_pretrained(model, checkpoint, subfolder=subfolder)77```78 79## Reproduction and evidence80 81The release includes the exact training JSONL and manifest plus four aggregate receipts and eight shard receipts for each reported evaluation row. Checksum files bind each evidence bundle.82 83These are assisted Fixed26 regression results using selected best-four cohorts, not pristine held-out benchmark claims. The Generalized C++ dataset, where applicable, explicitly overlaps six Fixed26 task IDs; consult its included manifest before comparing results.84 