Team Ai
Datasetpublic

foksly/orbit-mt-eval

ORBIT-MT-Eval: Expert Annotations and Prompts for Machine Translation Meta-Evaluation ORBIT-MT-Eval is an English-to-Russian reference dataset for machine translation meta-evaluation: comparing evaluator outputs against professional expert annotations. Each of the 600 source–translation segments has three independent expert annotations following the RATE protocol. In Beyond Prompts: A Systematic Study of LLM-Based Machine Translation Evaluators, two annotations serve as gold… See the full description on the dataset page: https://huggingface.co/datasets/foksly/orbit-mt-eval.

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
0likes22downloads
Dataset Card

ORBIT-MT-Eval: Expert Annotations and Prompts for Machine Translation Meta-Evaluation

ORBIT-MT-Eval is an English-to-Russian reference dataset for machine translation meta-evaluation: comparing evaluator outputs against professional expert annotations. Each of the 600 source–translation segments has three independent expert annotations following the RATE protocol. In Beyond Prompts: A Systematic Study of LLM-Based Machine Translation Evaluators, two annotations serve as gold references and the third is scored against each reference and averaged to produce the Human_E agreement baseline. The release also includes the prompts used by ORBIT and the locally maintained prompting baselines. GEMBA and unified-mqm-boosted-v5 are linked to their official upstream implementations.

Dataset at a glance

PropertyValue
Language directionEnglish → Russian
Source segments300
Source–translation segments600 (two translations for each source)
Professional expert annotators35
Expert annotations3 independent annotations per segment from the expert pool, for 1,800 records total
Meta-evaluation protocol2 reference annotations + 1 Human_E agreement baseline
Error spans9,771
Annotation schemeRATE severity 1–5 with separate Accuracy and Fluency scores
DomainsLiterary, news, social media, and speech

expert_1, expert_2, and expert_3 identify the three annotation slots within each segment. They are not person-level IDs and do not track individual annotators across the dataset.

Files

  • —`data/rate_annotations/test.jsonl`: the annotation records.
  • —`prompts/README.md`: readable prompt pages with copy-ready JSON messages objects.
  • —`prompts/prompts.jsonl`: the machine-readable prompt collection.

Load

python
from datasets import load_dataset

dataset = load_dataset("foksly/orbit-mt-eval", "rate_annotations", split="test")
print(dataset[0])

Paper results

EvaluatorRPF1R4+P4+MQM
Human_E46.844.345.565.365.375.3
Human_A-SbS39.347.142.855.763.268.0
Human_A-PW31.846.437.749.857.165.0
Human_WMT12.361.920.626.662.842.7
GEMBA-MQM24.643.031.345.051.058.7
GEMBA-ESA24.941.331.044.645.757.3
MQM-AE24.051.832.845.454.560.0
ESA-AE27.054.836.249.558.653.7
UMB-v527.844.534.248.756.457.3
xCOMET-XXL18.336.424.333.536.858.3
ORBIT-SC51.546.248.774.763.567.0
ORBIT-MC60.345.852.182.070.669.3

Values are on a 0–100 scale for the English→Russian evaluation set. R/P/F1 are micro-averaged span metrics excluding RATE severity 1, R4+/P4+ use severity ≥4, and MQM is pairwise ranking accuracy.

Prompting baselines and ORBIT-SC use GPT-5.4. ORBIT-MC fuses three evaluator backbones with an LLM fuser.

Citation

bibtex
@inproceedings{popov-etal-2026-beyond-prompts,
  title     = {Beyond Prompts: A Systematic Study of LLM-Based Machine Translation Evaluators},
  author    = {Popov, Dmitry and Bokhyan, Roman and Enikeeva, Ekaterina and
               Mekhraliev, Artem and Vysotsky, Stepan and Karpachev, Nikolay},
  booktitle = {Findings of the Association for Computational Linguistics: EMNLP 2026},
  year      = {2026}
}

License

The original RATE annotations and prompt material are released under the Apache License 2.0. Source and translation material retain their applicable upstream terms.