hafsteinn/bfcl-v3-multilingual
BFCL v3 in eight European languages Not manually verified yet. Every request is a machine translation. No native speaker has reviewed the set. Expect some unidiomatic phrasing and possibly a few meaning errors. Report problems in the Community tab. This is a translation of the five non-live, AST-checked categories of the Berkeley Function Calling Leaderboard v3 (simple 400, multiple 200, parallel 200, parallel_multiple 200, irrelevance 240; 1,240 items). The languages are… See the full description on the dataset page: https://huggingface.co/datasets/hafsteinn/bfcl-v3-multilingual.
BFCL v3 in eight European languages
Not manually verified yet. Every request is a machine translation. No native speaker has reviewed the set. Expect some unidiomatic phrasing and possibly a few meaning errors. Report problems in the Community tab.
This is a translation of the five non-live, AST-checked categories of the Berkeley Function Calling Leaderboard v3 (simple 400, multiple 200, parallel 200, parallel_multiple 200, irrelevance 240; 1,240 items). The languages are Icelandic, Faroese, Norwegian Bokmål, Norwegian Nynorsk, Swedish, Danish, Dutch and German. Icelandic has two independent translations, one by Claude Sonnet 5 and one by Claude Opus 5.5, so the effect of the translation system on scores can be measured.
Only the user request is translated. The function definitions and the ground truth stay in English, as in the original. The task therefore tests whether a model can map a request in language X onto English tool schemas. This is the usual situation with real APIs.
How it was made (28 September 2026)
- Translation. Via OpenRouter at temperature 0.2. The instructions: translate the request the way a native speaker would ask it, and do not add, drop, explain or answer anything.
- Protected literals. Each item has a list of strings that must survive verbatim:
- every number in the English request;
- every required ground-truth string argument that appears in the English request.
BFCL's checker compares argument values with the English ground truth, so a translated city name or product name would make a correct call fail. The list was given to the translator and checked afterwards, with one retry that names the missing strings.
- Items that still failed the check are marked
kept: falseand left out of the BFCL-format files. They are few; see the table. Some are false alarms: a short ground-truth value such as "LA" can match inside an unrelated English word ("Lastly"), and no faithful translation keeps it.
Known limitation: because of the protected literals, some English strings remain inside otherwise translated requests, for example dates such as "May 3rd 2022" and enum values such as "wild tiger". Scores measure tool use with these English fragments present, not fully native requests.
Files
translations.jsonl(the viewer'stestsplit):set,lang,translator,id(BFCL id),category,question(translated request),question_en,protected_literals,kept.bfcl/<set>/BFCL_v3_<category>.json: drop-in replacements for BFCL's own question files, with the kept items only and the user turn replaced.bfcl/possible_answer/BFCL_v3_<category>.json: BFCL's English ground truth, unchanged. Irrelevance has none, because the correct behaviour is to make no call. Score with BFCL's AST checker as usual; use the same ground truth for every language.
Licence and citation
BFCL is released under Apache-2.0, and so is this translation. Please cite the original:
@inproceedings{patil2025bfcl,
title={The Berkeley Function Calling Leaderboard (BFCL): From Tool Use to Agentic Evaluation of Large Language Models},
author={Patil, Shishir G. and Mao, Huanzhi and Yan, Fanjia and Ji, Charlie Cheng-Jie and Suresh, Vishnu and Stoica, Ion and Gonzalez, Joseph E.},
booktitle={Forty-second International Conference on Machine Learning},
year={2025}
}Produced by the TrustLLM SFT project at the University of Iceland.
