CompilingThings/compile-benchmark
CompilingThings Compile Benchmark for MQL5® This release evaluates compile success of generated MQL5 on two private held-out sets. The first is 300 Expert Advisor prompts, run on four arms: the base model, two tuned local models and one frontier API model. The second is 200 non-EA prompts (include files, custom indicators, scripts and services, 50 each), run on the three local arms, plus a stability re-run of one of them. The holdout results are attested, not fully verifiable:… See the full description on the dataset page: https://huggingface.co/datasets/CompilingThings/compile-benchmark.
v1.1.0: pin the LFS attribute for corpus_row_hashes.json in .gitattributes and re-seal its manifest line
CompilingThings Compile Benchmark for MQL5 - v1.1.0
LICENSE: remove the hosting-platform savings clause; SHA256SUMS: recompute LICENSE and README.md entries; verify_public_release.py ALL CHECKS PASSED on this tree
Dataset card: enable Viewer for the two public benchmark files (prompts, per_item_results); add size_categories n<1K; body unchanged
Dataset card tags: keep mql5, drop metatrader-5 and code (trademark-conservative); body unchanged
Add dataset-card discovery metadata (pretty_name, language, task_categories, tags); body unchanged
CompilingThings Compile Benchmark for MQL5 - v1.0.0
initial commit
