Team Ai
Datasetpublic

TokenBender/glm47-pie-cpp-posttraining-data

GLM-4.7-Flash PIE C++ Post-Training Data The exact prepared dataset used for the GLM-4.7-Flash C++ performance post-training runs. Splits File Rows Purpose sft/train.jsonl 7,864 Supervised fine-tuning grpo/train.jsonl 7,887 GRPO prompt and reward evaluation eval/validation.jsonl 1,259 Full held-out evaluation eval/validation_mini126.jsonl 126 Fast evaluation eval/validation_mini4.jsonl 4 Smoke evaluation tasks.tar.gz 9,146 task JSONs Reward… See the full description on the dataset page: https://huggingface.co/datasets/TokenBender/glm47-pie-cpp-posttraining-data.

sourceHugging Faceotherupdated 3mo agoView on Hugging Face
0likes42downloads
README.md68 linesDownload Raw Back to root
1---2license: other3task_categories:4  - text-generation5language:6  - en7tags:8  - code9  - code-optimization10  - cpp11  - sft12  - grpo13configs:14  - config_name: sft15    data_files:16      - split: train17        path: sft/train.jsonl18  - config_name: grpo19    data_files:20      - split: train21        path: grpo/train.jsonl22  - config_name: evaluation23    data_files:24      - split: validation25        path: eval/validation.jsonl26---27 28# GLM-4.7-Flash PIE C++ Post-Training Data29 30The exact prepared dataset used for the GLM-4.7-Flash C++ performance31post-training runs.32 33## Splits34 35| File | Rows | Purpose |36| --- | ---: | --- |37| `sft/train.jsonl` | 7,864 | Supervised fine-tuning |38| `grpo/train.jsonl` | 7,887 | GRPO prompt and reward evaluation |39| `eval/validation.jsonl` | 1,259 | Full held-out evaluation |40| `eval/validation_mini126.jsonl` | 126 | Fast evaluation |41| `eval/validation_mini4.jsonl` | 4 | Smoke evaluation |42| `tasks.tar.gz` | 9,146 task JSONs | Reward scoring and evaluation harness inputs |43 44Each row carries a stable task ID, problem ID, split, prompt, label, and task45metadata. SFT rows additionally contain the user and assistant messages used by46the trainer. The task archive contains every `tasks/...` path referenced by the47training and evaluation rows. The canonical repository downloader verifies and48extracts it into `data/tasks`.49 50## Preparation51 52The source tasks were filtered by compiling and executing the oracle answer in53the C++ sandbox. Training rows were retained when the response format was valid,54all tests passed, and the reward result was `correct`.55 56The source preparation run scored 8,350 training tasks, retained 7,887, and57rejected 463. The complete preparation receipt is in `manifest.json`.58 59## Provenance60 61This is a prepared derivative of the62[PIE C++ performance dataset](https://github.com/madaan/pie-perf), which is63based on IBM Project CodeNet. Use and redistribution are subject to the64applicable upstream dataset terms.65 66Training code:67[TokenBender/browser-is-all-you-need](https://github.com/tokenbender/browser-is-all-you-need/tree/client/glm47-h100-posttraining)68