Team Ai
Datasetpublic

LaurelWings/rcga-evaluation-data

RCGA / LoopSFT evaluation input snapshots Exact benchmark inputs used in our Qwen3-1.7B, Qwen3-4B and Qwen3-30B-A3B experiments. These are evaluation questions, prompts, labels and tests, not model-generated answers or reported scores. Already frozen subsets are stored directly, including original row order and prompt formatting. No new random split was made during backup. Collections data/evalbench/: all 31 historical prompt variants. The active standard14… See the full description on the dataset page: https://huggingface.co/datasets/LaurelWings/rcga-evaluation-data.

sourceHugging Faceotherupdated 10d agoView on Hugging Face
0likes134downloads
Dataset Card

RCGA / LoopSFT evaluation input snapshots

Exact benchmark inputs used in our Qwen3-1.7B, Qwen3-4B and Qwen3-30B-A3B experiments. These are evaluation questions, prompts, labels and tests, not model-generated answers or reported scores. Already frozen subsets are stored directly, including original row order and prompt formatting. No new random split was made during backup.

Collections

  • —data/evalbench/: all 31 historical prompt variants. The active standard14 protocol uses 13 distinct files; MATH500 and MATH500 Flex score the same 250 prompts differently. This set was used by the 1.7B/30B standard benchmark runs. Coding16 and AIME multi-sample runs reuse these same input rows; repetition counts and decoding settings are separate from the data split.
  • —data/qwen4b_frozen/: the distinct historical 4B frozen input suite, with its original manifest. Do not substitute the standard14 subsets for these files.
  • —data/coda_probe/samples.jsonl: 128 fixed likelihood-probe examples, with target boundaries and token IDs. These were not deduplicated against training data and are not established held-out examples.
  • —auxiliary/: historical EvalPlus task releases used by our own frozen evaluation and reference-probe data.

See inventory.csv for every file and row count, SPLITS.md for sampling/prompt rules, and SOURCES.md for upstream attribution. Each JSONL file is a separate Hugging Face config to avoid mixing incompatible schemas.

python
from datasets import load_dataset
x = load_dataset('LaurelWings/rcga-evaluation-data', 'evalbench_math500_fewshot', split='test')

For exact original JSONL bytes rather than Arrow conversion:

python
from huggingface_hub import snapshot_download
snapshot_download('LaurelWings/rcga-evaluation-data', repo_type='dataset', local_dir='evaluation-backup')

protocols/standard14_original.yaml preserves historical absolute paths. standard14_relative.yaml points to this repository layout; run from the repository root or resolve these paths explicitly in your launcher. Raw source recipes are archived for provenance, and some retain old /root or /workspace paths; use the frozen data files for exact recovery. No GPU training or evaluation is launched by downloading this repository.