CompilingThings/compile-benchmark
CompilingThings Compile Benchmark for MQL5® This release evaluates compile success of generated MQL5 on two private held-out sets. The first is 300 Expert Advisor prompts, run on four arms: the base model, two tuned local models and one frontier API model. The second is 200 non-EA prompts (include files, custom indicators, scripts and services, 50 each), run on the three local arms, plus a stability re-run of one of them. The holdout results are attested, not fully verifiable:… See the full description on the dataset page: https://huggingface.co/datasets/CompilingThings/compile-benchmark.
1131
1cff-version: 1.2.02message: If you use this benchmark, cite it as below.3title: CompilingThings Compile Benchmark for MQL5®4type: dataset5authors:6 - name: CompilingThings7version: 1.1.08date-released: 2026-09-129identifiers:10 - type: other11 value: CompilingThings/compile-benchmark-v1.1.012 description: Release identifier.13abstract: >-14 A paired benchmark of whether generated MQL5® source compiles. Version 1.1.015 carries the 184 public prompts and their 1.0.0 per-item results unchanged.16 It adds hash-keyed per-item results on two private held-out sets: 300 Expert17 Advisor prompts run on four arms, and 200 non-EA prompts run on three local18 arms plus a stability re-run. It also adds a Q8_0 bridge re-run of the 18419 public prompts, the serving template and system prompts, and a corpus20 row-hash manifest.21 Released under the CompilingThings Benchmark Evaluation Licence v1.0; see the LICENSE file for terms.22 MQL5® and MetaTrader 5® are registered trademarks of MetaQuotes Ltd.23 CompilingThings is an independent project. No affiliation, sponsorship, certification, endorsement, or approval by MetaQuotes Ltd. is claimed.24 