Team Ai
Datasetpublic

CompilingThings/compile-benchmark

CompilingThings Compile Benchmark for MQL5® This release evaluates compile success of generated MQL5 on two private held-out sets. The first is 300 Expert Advisor prompts, run on four arms: the base model, two tuned local models and one frontier API model. The second is 200 non-EA prompts (include files, custom indicators, scripts and services, 50 each), run on the three local arms, plus a stability re-run of one of them. The holdout results are attested, not fully verifiable:… See the full description on the dataset page: https://huggingface.co/datasets/CompilingThings/compile-benchmark.

sourceHugging Faceotherupdated 27d agoView on Hugging Face
1likes131downloads
CITATION.cff24 linesDownload Raw Back to root
1cff-version: 1.2.02message: If you use this benchmark, cite it as below.3title: CompilingThings Compile Benchmark for MQL5®4type: dataset5authors:6  - name: CompilingThings7version: 1.1.08date-released: 2026-09-129identifiers:10  - type: other11    value: CompilingThings/compile-benchmark-v1.1.012    description: Release identifier.13abstract: >-14  A paired benchmark of whether generated MQL5® source compiles. Version 1.1.015  carries the 184 public prompts and their 1.0.0 per-item results unchanged.16  It adds hash-keyed per-item results on two private held-out sets: 300 Expert17  Advisor prompts run on four arms, and 200 non-EA prompts run on three local18  arms plus a stability re-run. It also adds a Q8_0 bridge re-run of the 18419  public prompts, the serving template and system prompts, and a corpus20  row-hash manifest.21  Released under the CompilingThings Benchmark Evaluation Licence v1.0; see the LICENSE file for terms.22  MQL5® and MetaTrader 5® are registered trademarks of MetaQuotes Ltd.23  CompilingThings is an independent project. No affiliation, sponsorship, certification, endorsement, or approval by MetaQuotes Ltd. is claimed.24