jumelet/multiblimp
MultiBLiMP MultiBLiMP is a massively Multilingual Benchmark for Linguistic Minimal Pairs. The dataset is composed of synthetic pairs generated using Universal Dependencies and UniMorph. The paper can be found here. We split the data set by language: each language consists of a single .tsv file. The rows contain many attributes for a particular pair, most important are the sen and wrong_sen fields, which we use for evaluating the language models. Using MultiBLiMP… See the full description on the dataset page: https://huggingface.co/datasets/jumelet/multiblimp.
This repository belongs to jumelet on Hugging Face.
Team Ai never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
