AmazonScience/mxeval
A collection of execution-based multi-lingual benchmark for code generation.
046
1---2dataset_info:3 features:4 - name: task_id5 dtype: string6 - name: language7 dtype: string8 - name: prompt9 dtype: string10 - name: test11 dtype: string12 - name: entry_point13 dtype: string14 splits:15 - name: multilingual-humaneval_python16 num_bytes: 16571617 num_examples: 16418 download_size: 6798319 dataset_size: 16571620license: apache-2.021task_categories:22- text-generation23tags:24- mxeval25- code-generation26- mbxp27- multi-humaneval28- mathqax29pretty_name: mxeval30language:31- en32---33# MxEval34**M**ultilingual E**x**ecution **Eval**uation35 36## Table of Contents37- [MxEval](#MxEval)38 - [Table of Contents](#table-of-contents)39 - [Dataset Description](#dataset-description)40 - [Dataset Summary](#dataset-summary)41 - [Supported Tasks and Leaderboards](#supported-tasks-and-leaderboards)42 - [Languages](#languages)43 - [Dataset Structure](#dataset-structure)44 - [Data Instances](#data-instances)45 - [Data Fields](#data-fields)46 - [Data Splits](#data-splits)47 - [Dataset Creation](#dataset-creation)48 - [Curation Rationale](#curation-rationale)49 - [Personal and Sensitive Information](#personal-and-sensitive-information)50 - [Social Impact of Dataset](#social-impact-of-dataset)51 - [Executional Correctness](#execution)52 - [Execution Example](#execution-example)53 - [Considerations for Using the Data](#considerations-for-using-the-data)54 - [Additional Information](#additional-information)55 - [Dataset Curators](#dataset-curators)56 - [Licensing Information](#licensing-information)57 - [Citation Information](#citation-information)58 - [Contributions](#contributions)59 60## Dataset Description61 62- **Repository:** [GitHub Repository](https://github.com/amazon-science/mxeval)63- **Paper:** [Multi-lingual Evaluation of Code Generation Models](https://openreview.net/forum?id=Bo7eeXm6An8)64 65### Dataset Summary66 67This repository contains data and code to perform execution-based multi-lingual evaluation of code generation capabilities and the corresponding data,68namely, a multi-lingual benchmark MBXP, multi-lingual MathQA and multi-lingual HumanEval.69<br>Results and findings can be found in the paper ["Multi-lingual Evaluation of Code Generation Models"](https://arxiv.org/abs/2210.14868).70 71 72### Supported Tasks and Leaderboards73* [MBXP](https://huggingface.co/datasets/mxeval/mbxp)74* [Multi-HumanEval](https://huggingface.co/datasets/mxeval/multi-humaneval)75* [MathQA-X](https://huggingface.co/datasets/mxeval/mathqa-x)76 77### Languages78The programming problems are written in multiple programming languages and contain English natural text in comments and docstrings.79 80 81## Dataset Structure82To lookup currently supported datasets83```python84get_dataset_config_names("AmazonScience/mxeval")85['mathqa-x', 'mbxp', 'multi-humaneval']86```87To load a specific dataset and language88```python89from datasets import load_dataset90load_dataset("AmazonScience/mxeval", "mbxp", split="python")91Dataset({92 features: ['task_id', 'language', 'prompt', 'test', 'entry_point', 'description', 'canonical_solution'],93 num_rows: 97494})95```96 97### Data Instances98 99An example of a dataset instance:100 101```python102{103 "task_id": "MBSCP/6",104 "language": "scala",105 "prompt": "object Main extends App {\n /**\n * You are an expert Scala programmer, and here is your task.\n * * Write a Scala function to check whether the two numbers differ at one bit position only or not.\n *\n * >>> differAtOneBitPos(13, 9)\n * true\n * >>> differAtOneBitPos(15, 8)\n * false\n * >>> differAtOneBitPos(2, 4)\n * false\n */\n def differAtOneBitPos(a : Int, b : Int) : Boolean = {\n",106 "test": "\n\n var arg00 : Int = 13\n var arg01 : Int = 9\n var x0 : Boolean = differAtOneBitPos(arg00, arg01)\n var v0 : Boolean = true\n assert(x0 == v0, \"Exception -- test case 0 did not pass. x0 = \" + x0)\n\n var arg10 : Int = 15\n var arg11 : Int = 8\n var x1 : Boolean = differAtOneBitPos(arg10, arg11)\n var v1 : Boolean = false\n assert(x1 == v1, \"Exception -- test case 1 did not pass. x1 = \" + x1)\n\n var arg20 : Int = 2\n var arg21 : Int = 4\n var x2 : Boolean = differAtOneBitPos(arg20, arg21)\n var v2 : Boolean = false\n assert(x2 == v2, \"Exception -- test case 2 did not pass. x2 = \" + x2)\n\n\n}\n",107 "entry_point": "differAtOneBitPos",108 "description": "Write a Scala function to check whether the two numbers differ at one bit position only or not."109}110```111 112### Data Fields113 114- `task_id`: identifier for the data sample115- `prompt`: input for the model containing function header and docstrings116- `canonical_solution`: solution for the problem in the `prompt`117- `description`: task description118- `test`: contains function to test generated code for correctness119- `entry_point`: entry point for test120- `language`: programming lanuage identifier to call the appropriate subprocess call for program execution121 122 123### Data Splits124 125 - HumanXEval126 - Python 127 - Java 128 - JavaScript129 - Csharp130 - CPP131 - Go132 - Kotlin133 - PHP134 - Perl135 - Ruby136 - Swift137 - Scala138 - MBXP139 - Python140 - Java141 - JavaScript142 - TypeScript143 - Csharp144 - CPP145 - Go146 - Kotlin147 - PHP148 - Perl149 - Ruby150 - Swift151 - Scala152 - MathQA153 - Python154 - Java155 - JavaScript156 157 158## Dataset Creation159 160### Curation Rationale161 162Since code generation models are often trained on dumps of GitHub a dataset not included in the dump was necessary to properly evaluate the model. However, since this dataset was published on GitHub it is likely to be included in future dumps.163 164### Personal and Sensitive Information165 166None.167 168### Social Impact of Dataset169With this dataset code generating models can be better evaluated which leads to fewer issues introduced when using such models.170 171### Dataset Curators172AWS AI Labs173 174## Execution175 176### Execution Example177Install the repo [mxeval](https://github.com/amazon-science/mxeval) to execute generations or canonical solutions for the prompts from this dataset.178 179```python180>>> from datasets import load_dataset181>>> from mxeval.execution import check_correctness182>>> mbxp_python = load_dataset("AmazonScience/mxeval", "mbxp", split="python")183>>> example_problem = mbxp_python[0]184>>> check_correctness(example_problem, example_problem["canonical_solution"], timeout=20.0)185{'task_id': 'MBPP/1', 'passed': True, 'result': 'passed', 'completion_id': None, 'time_elapsed': 10.582208633422852}186```187### Considerations for Using the Data188Make sure to sandbox the execution environment since generated code samples can be harmful.189 190 191### Licensing Information192 193[LICENSE](https://huggingface.co/datasets/AmazonScience/mxeval/blob/main/LICENSE) <br>194[THIRD PARTY LICENSES](https://huggingface.co/datasets/AmazonScience/mxeval/blob/main/THIRD_PARTY_LICENSES)195 196# Citation Information197```198@article{mbxp_athiwaratkun2022,199 title = {Multi-lingual Evaluation of Code Generation Models},200 author = {Athiwaratkun, Ben and201 Gouda, Sanjay Krishna and202 Wang, Zijian and203 Li, Xiaopeng and204 Tian, Yuchen and205 Tan, Ming206 and Ahmad, Wasi Uddin and207 Wang, Shiqi and208 Sun, Qing and209 Shang, Mingyue and210 Gonugondla, Sujan Kumar and211 Ding, Hantian and212 Kumar, Varun and213 Fulton, Nathan and214 Farahani, Arash and215 Jain, Siddhartha and216 Giaquinto, Robert and217 Qian, Haifeng and218 Ramanathan, Murali Krishna and219 Nallapati, Ramesh and220 Ray, Baishakhi and221 Bhatia, Parminder and222 Sengupta, Sudipta and223 Roth, Dan and224 Xiang, Bing},225 doi = {10.48550/ARXIV.2210.14868},226 url = {https://arxiv.org/abs/2210.14868},227 keywords = {Machine Learning (cs.LG), Computation and Language (cs.CL), FOS: Computer and information sciences, FOS: Computer and information sciences},228 publisher = {arXiv},229 year = {2022},230 copyright = {Creative Commons Attribution 4.0 International}231}232```233 234# Contributions235 236[skgouda@](https://github.com/sk-g) [benathi@](https://github.com/benathi)