foksly/orbit-mt-eval
ORBIT-MT-Eval: Expert Annotations and Prompts for Machine Translation Meta-Evaluation ORBIT-MT-Eval is an English-to-Russian reference dataset for machine translation meta-evaluation: comparing evaluator outputs against professional expert annotations. Each of the 600 source–translation segments has three independent expert annotations following the RATE protocol. In Beyond Prompts: A Systematic Study of LLM-Based Machine Translation Evaluators, two annotations serve as gold… See the full description on the dataset page: https://huggingface.co/datasets/foksly/orbit-mt-eval.
ORBIT-MT-Eval: Expert Annotations and Prompts for Machine Translation Meta-Evaluation
ORBIT-MT-Eval is an English-to-Russian reference dataset for machine translation meta-evaluation: comparing evaluator outputs against professional expert annotations. Each of the 600 source–translation segments has three independent expert annotations following the RATE protocol. In Beyond Prompts: A Systematic Study of LLM-Based Machine Translation Evaluators, two annotations serve as gold references and the third is scored against each reference and averaged to produce the Human_E agreement baseline. The release also includes the prompts used by ORBIT and the locally maintained prompting baselines. GEMBA and unified-mqm-boosted-v5 are linked to their official upstream implementations.
Dataset at a glance
expert_1, expert_2, and expert_3 identify the three annotation slots within each segment. They are not person-level IDs and do not track individual annotators across the dataset.
Files
- `data/rate_annotations/test.jsonl`: the annotation records.
- `prompts/README.md`: readable prompt pages with copy-ready JSON
messagesobjects. - `prompts/prompts.jsonl`: the machine-readable prompt collection.
Load
from datasets import load_dataset
dataset = load_dataset("foksly/orbit-mt-eval", "rate_annotations", split="test")
print(dataset[0])Paper results
Values are on a 0–100 scale for the English→Russian evaluation set. R/P/F1 are micro-averaged span metrics excluding RATE severity 1, R4+/P4+ use severity ≥4, and MQM is pairwise ranking accuracy.
Prompting baselines and ORBIT-SC use GPT-5.4. ORBIT-MC fuses three evaluator backbones with an LLM fuser.
Citation
@inproceedings{popov-etal-2026-beyond-prompts,
title = {Beyond Prompts: A Systematic Study of LLM-Based Machine Translation Evaluators},
author = {Popov, Dmitry and Bokhyan, Roman and Enikeeva, Ekaterina and
Mekhraliev, Artem and Vysotsky, Stepan and Karpachev, Nikolay},
booktitle = {Findings of the Association for Computational Linguistics: EMNLP 2026},
year = {2026}
}License
The original RATE annotations and prompt material are released under the Apache License 2.0. Source and translation material retain their applicable upstream terms.
