Team Ai
Datasetpublic

jumelet/multiblimp

MultiBLiMP MultiBLiMP is a massively Multilingual Benchmark for Linguistic Minimal Pairs. The dataset is composed of synthetic pairs generated using Universal Dependencies and UniMorph. The paper can be found here. We split the data set by language: each language consists of a single .tsv file. The rows contain many attributes for a particular pair, most important are the sen and wrong_sen fields, which we use for evaluating the language models. Using MultiBLiMP… See the full description on the dataset page: https://huggingface.co/datasets/jumelet/multiblimp.

sourceHugging Facecc-by-4.0updated 1y agoView on Hugging Face
17likes9.8kdownloads
Dataset Card

MultiBLiMP

MultiBLiMP is a massively Multilingual Benchmark for Linguistic Minimal Pairs. The dataset is composed of synthetic pairs generated using Universal Dependencies and UniMorph.

The paper can be found here.

We split the data set by language: each language consists of a single .tsv file. The rows contain many attributes for a particular pair, most important are the sen and wrong_sen fields, which we use for evaluating the language models.

Using MultiBLiMP

To download the final datasets for a given language, you can usethe example code. Make sure to use the ISO 639-3 code, e.g. eng for English (see table below).

from huggingface_hub import hf_hub_download

hf_hub_download(repo_id="jumelet/multiblimp", filename=f"eng/data.tsv", repo_type="dataset", local_dir='hf_cache/')

To run MultiBLiMP on a model, follow the instructions in this repo. Example evaluation code:

python scripts/lm_eval/eval_model.py 
        --model meta-llama/Meta-Llama-3-8B 
        --data_dir hf_cache/eng/ 
        --src_dir multiblimp 
        --results_dir multiblimp_results/ 
        --cache_dir hf_cache

Languages

This table contains the languages covered in MultiBLiMP and the number of items for each language.

ISO CodeLanguagen
abkAbkhazian40
aqzAkuntsu14
sqiAlbanian243
amhAmharic112
grcAncient Greek3695
hboAncient Hebrew983
apuApurinã28
hyeArmenian1415
eusBasque273
belBelarusian2570
benBengali21
bhoBhojpuri34
borBorôro241
breBreton260
bulBulgarian2458
buaBuriat103
catCatalan2284
chuChurch Slavonic4166
xclClassical Armenian1623
cesCzech4256
danDanish50
nldDutch2331
egyEgyptian (Ancient)22
engEnglish770
myvErzya464
estEstonian2575
faoFaroese232
finFinnish2570
fraFrench2548
glgGalician753
katGeorgian204
deuGerman2298
alnGheg Albanian677
gotGothic1579
gujGujarati7
hebHebrew2330
azzH-P Nahuatl207
hinHindi1447
hitHittite50
hunHungarian845
islIcelandic2801
gleIrish28
itaItalian2999
qucK'iche'131
xnrKangri86
krlKarelian260
kxhKaro (Ethiopia)120
kazKazakh173
kirKirghiz185
koiKomi-Permyak43
kpvKomi-Zyrian320
latLatin3149
lavLatvian3032
lijLigurian254
litLithuanian1180
oloLivvi190
ndsLow German1774
mkdMacedonian39
marMarathi460
frmMiddle French294
ellModern Greek1096
mdfMoksha82
yrlNhengatu720
pcmNigerian Pidgin26
kmrNorthern Kurdish544
smeNorthern Sami2536
froOld French1976
orvOld Russian4615
otaOttoman Turkish99
fasPersian2553
xpgPhrygian50
polPolish3272
porPortuguese3048
ronRomanian2056
rusRussian3832
sanSanskrit4442
glaScottish Gaelic66
hbsSerbo-Croatian3286
smsSkolt Sami263
slkSlovak4145
slvSlovenian4483
spaSpanish2541
arbStandard Arabic1215
sweSwedish201
tamTamil382
ttcTektiteko69
tpnTupinambá9
turTurkish1742
uigUighur758
ukrUkrainian2744
hsbUpper Sorbian186
urdUrdu550
urbUrubú-Kaapor13
uzbUzbek50
vepVeps187
wbpWarlpiri12
cymWelsh1120
hywWestern Armenian1153
wolWolof705
sahYakut144
nhiTenango Nahuatl38

Citation Information

@misc{jumelet2025multiblimp10massivelymultilingual,
      title={MultiBLiMP 1.0: A Massively Multilingual Benchmark of Linguistic Minimal Pairs}, 
      author={Jaap Jumelet and Leonie Weissweiler and Arianna Bisazza},
      year={2025},
      eprint={2504.02768},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2504.02768}, 
}