Team Ai
Datasetpublic

AmazonScience/massive-agents

sourceHugging Facecc-by-4.0updated 5mo agoView on Hugging Face
2likes374downloads
CITATION.cff61 linesDownload Raw Back to root
1cff-version: 1.2.02message: "If you use this work, please cite it as below."3title: "MASSIVE-Agents: A Benchmark for Multilingual Function-Calling in 52 Languages"4type: conference-paper5authors:6  - family-names: "Kulkarni"7    given-names: "Mayank"8  - family-names: "Mazzia"9    given-names: "Vittorio"10  - family-names: "Gaspers"11    given-names: "Judith"12  - family-names: "Hench"13    given-names: "Chris"14  - family-names: "FitzGerald"15    given-names: "Jack"16year: 202517month: 1118conference:19  name: "Findings of the Association for Computational Linguistics: EMNLP 2025"20  location: "Suzhou, China"21publisher:22  name: "Association for Computational Linguistics"23pages: "20193-20215"24doi: "10.18653/v1/2025.findings-emnlp.1099"25isbn: "979-8-89176-335-7"26url: "https://aclanthology.org/2025.findings-emnlp.1099/"27abstract: >28  We present MASSIVE-Agents, a new benchmark for assessing multilingual29  function calling across 52 languages. We created MASSIVE-Agents by30  cleaning the original MASSIVE dataset and then reformatting it for31  evaluation within the Berkeley Function-Calling Leaderboard (BFCL)32  framework. The full benchmark comprises 47,020 samples with an average33  of 904 samples per language, covering 55 different functions and 28634  arguments. We benchmarked 21 models using Amazon Bedrock and present35  the results along with associated analyses. MASSIVE-Agents is36  challenging, with the top model Nova Premier achieving an average37  Abstract Syntax Tree (AST) Accuracy of 34.05% across all languages,38  with performance varying significantly from 57.37% for English to as39  low as 6.81% for Amharic. Some models, particularly smaller ones,40  yielded a score of zero for the more difficult languages. Additionally,41  we provide results from ablations using a custom 1-shot prompt,42  ablations with prompts translated into different languages, and43  comparisons based on model latency.44preferred-citation:45  type: paper-conference46  authors:47    - family-names: "Kulkarni"48      given-names: "Mayank"49    - family-names: "Mazzia"50      given-names: "Vittorio"51    - family-names: "Gaspers"52      given-names: "Judith"53    - family-names: "Hench"54      given-names: "Chris"55    - family-names: "FitzGerald"56      given-names: "Jack"57  title: "MASSIVE-Agents: A Benchmark for Multilingual Function-Calling in 52 Languages"58  year: 202559  conference:60    name: "Findings of the Association for Computational Linguistics: EMNLP 2025"61  doi: "10.18653/v1/2025.findings-emnlp.1099"