Polyglot-or-Not/Fact-Completion
Dataset Card Homepage: https://bit.ly/ischool-berkeley-capstone Repository: https://github.com/daniel-furman/Capstone Point of Contact: daniel_furman@berkeley.edu Dataset Summary This is the dataset for Polyglot or Not?: Measuring Multilingual Encyclopedic Knowledge Retrieval from Foundation Language Models. Test Description Given a factual association such as The capital of France is Paris, we determine whether a model adequately "knows" this… See the full description on the dataset page: https://huggingface.co/datasets/Polyglot-or-Not/Fact-Completion.
Dataset Card
- Homepage: https://bit.ly/ischool-berkeley-capstone
- Repository: https://github.com/daniel-furman/Capstone
- Point of Contact: daniel_furman@berkeley.edu
Dataset Summary
This is the dataset for Polyglot or Not?: Measuring Multilingual Encyclopedic Knowledge Retrieval from Foundation Language Models.
Test Description
Given a factual association such as The capital of France is Paris**, we determine whether a model adequately "knows" this information with the following test:
- Step 1: prompt the model to predict the likelihood of the token Paris following The Capital of France is
- Step 2: prompt the model to predict the average likelihood of a set of false, counterfactual tokens following the same stem.
If the value from 1 is greater than the value from 2 we conclude that model adequately recalls that fact. Formally, this is an application of the Contrastive Knowledge Assessment proposed in [[1][bib]].
For every foundation model of interest (like LLaMA), we perform this assessment on a set of facts translated into 20 languages. All told, we score foundation models on 303k fact-completions (results).
We also score monolingual models (like GPT-2) on English-only fact-completion (results).
Languages
The dataset covers 20 languages, which use either the Latin or Cyrillic scripts: bg, ca, cs, da, de, en, es, fr, hr, hu, it, nl, pl, pt, ro, ru, sl, sr, sv, uk.
Data Splits
The dataset splits correspond to the 20 languages above.
Source Data
We sourced the English cut of the dataset from [1] and [2] and used the Google Translate API to produce the other 19 language cuts.
Licensing Information
The dataset is licensed under the Apache 2.0 license and may be used with the corresponding affordances without limit.
Citation Information
@misc{schott2023polyglot,
doi = {10.48550/arXiv.2305.13675},
title={Polyglot or Not? Measuring Multilingual Encyclopedic Knowledge Retrieval from Foundation Language Models},
author={Tim Schott and Daniel Furman and Shreshta Bhat},
year={2023},
eprint={2305.13675,
archivePrefix={arXiv},
primaryClass={cs.CL}
}Bibliography
[1] Dong, Qingxiu, Damai Dai, Yifan Song, Jingjing Xu, Zhifang Sui, and Lei Li. "Calibrating Factual Knowledge in Pretrained Language Models". In Findings of the Association for Computational Linguistics: EMNLP 2022. [arXiv:2210.03329][cka] (2022).
@misc{dong2022calibrating,
doi = {10.48550/arXiv.2210.03329},
title={Calibrating Factual Knowledge in Pretrained Language Models},
author={Qingxiu Dong and Damai Dai and Yifan Song and Jingjing Xu and Zhifang Sui and Lei Li},
year={2022},
eprint={2210.03329},
archivePrefix={arXiv},
primaryClass={cs.CL}
}[2] Meng, Kevin, Arnab Sen Sharma, Alex Andonian, Yonatan Belinkov, and David Bau. "Mass Editing Memory in a Transformer." arXiv preprint [arXiv:2210.07229][memit] (2022).
@misc{meng2022massediting,
doi = {10.48550/arXiv.2210.07229},
title={Mass-Editing Memory in a Transformer},
author={Kevin Meng and Arnab Sen Sharma and Alex Andonian and Yonatan Belinkov and David Bau},
year={2022},
eprint={2210.07229},
archivePrefix={arXiv},
primaryClass={cs.CL}
}