Team Ai
Datasetpublic

AmazonScience/mxeval

A collection of execution-based multi-lingual benchmark for code generation.

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
0likes46downloads
README.md236 linesDownload Raw Back to root
1---2dataset_info:3  features:4  - name: task_id5    dtype: string6  - name: language7    dtype: string8  - name: prompt9    dtype: string10  - name: test11    dtype: string12  - name: entry_point13    dtype: string14  splits:15  - name: multilingual-humaneval_python16    num_bytes: 16571617    num_examples: 16418  download_size: 6798319  dataset_size: 16571620license: apache-2.021task_categories:22- text-generation23tags:24- mxeval25- code-generation26- mbxp27- multi-humaneval28- mathqax29pretty_name: mxeval30language:31- en32---33# MxEval34**M**ultilingual E**x**ecution **Eval**uation35 36## Table of Contents37- [MxEval](#MxEval)38  - [Table of Contents](#table-of-contents)39  - [Dataset Description](#dataset-description)40    - [Dataset Summary](#dataset-summary)41    - [Supported Tasks and Leaderboards](#supported-tasks-and-leaderboards)42    - [Languages](#languages)43  - [Dataset Structure](#dataset-structure)44    - [Data Instances](#data-instances)45    - [Data Fields](#data-fields)46    - [Data Splits](#data-splits)47  - [Dataset Creation](#dataset-creation)48    - [Curation Rationale](#curation-rationale)49    - [Personal and Sensitive Information](#personal-and-sensitive-information)50    - [Social Impact of Dataset](#social-impact-of-dataset)51  - [Executional Correctness](#execution)52    - [Execution Example](#execution-example)53    - [Considerations for Using the Data](#considerations-for-using-the-data)54  - [Additional Information](#additional-information)55    - [Dataset Curators](#dataset-curators)56    - [Licensing Information](#licensing-information)57    - [Citation Information](#citation-information)58    - [Contributions](#contributions)59 60## Dataset Description61 62- **Repository:** [GitHub Repository](https://github.com/amazon-science/mxeval)63- **Paper:** [Multi-lingual Evaluation of Code Generation Models](https://openreview.net/forum?id=Bo7eeXm6An8)64 65### Dataset Summary66 67This repository contains data and code to perform execution-based multi-lingual evaluation of code generation capabilities and the corresponding data,68namely, a multi-lingual benchmark MBXP, multi-lingual MathQA and multi-lingual HumanEval.69<br>Results and findings can be found in the paper ["Multi-lingual Evaluation of Code Generation Models"](https://arxiv.org/abs/2210.14868).70 71 72### Supported Tasks and Leaderboards73* [MBXP](https://huggingface.co/datasets/mxeval/mbxp)74* [Multi-HumanEval](https://huggingface.co/datasets/mxeval/multi-humaneval)75* [MathQA-X](https://huggingface.co/datasets/mxeval/mathqa-x)76 77### Languages78The programming problems are written in multiple programming languages and contain English natural text in comments and docstrings.79 80 81## Dataset Structure82To lookup currently supported datasets83```python84get_dataset_config_names("AmazonScience/mxeval")85['mathqa-x', 'mbxp', 'multi-humaneval']86```87To load a specific dataset and language88```python89from datasets import load_dataset90load_dataset("AmazonScience/mxeval", "mbxp", split="python")91Dataset({92    features: ['task_id', 'language', 'prompt', 'test', 'entry_point', 'description', 'canonical_solution'],93    num_rows: 97494})95```96 97### Data Instances98 99An example of a dataset instance:100 101```python102{103  "task_id": "MBSCP/6",104  "language": "scala",105  "prompt": "object Main extends App {\n    /**\n     * You are an expert Scala programmer, and here is your task.\n     * * Write a Scala function to check whether the two numbers differ at one bit position only or not.\n     *\n     * >>> differAtOneBitPos(13, 9)\n     * true\n     * >>> differAtOneBitPos(15, 8)\n     * false\n     * >>> differAtOneBitPos(2, 4)\n     * false\n     */\n    def differAtOneBitPos(a : Int, b : Int) : Boolean = {\n",106  "test": "\n\n    var arg00 : Int = 13\n    var arg01 : Int = 9\n    var x0 : Boolean = differAtOneBitPos(arg00, arg01)\n    var v0 : Boolean = true\n    assert(x0 == v0, \"Exception -- test case 0 did not pass. x0 = \" + x0)\n\n    var arg10 : Int = 15\n    var arg11 : Int = 8\n    var x1 : Boolean = differAtOneBitPos(arg10, arg11)\n    var v1 : Boolean = false\n    assert(x1 == v1, \"Exception -- test case 1 did not pass. x1 = \" + x1)\n\n    var arg20 : Int = 2\n    var arg21 : Int = 4\n    var x2 : Boolean = differAtOneBitPos(arg20, arg21)\n    var v2 : Boolean = false\n    assert(x2 == v2, \"Exception -- test case 2 did not pass. x2 = \" + x2)\n\n\n}\n",107  "entry_point": "differAtOneBitPos",108  "description": "Write a Scala function to check whether the two numbers differ at one bit position only or not."109}110```111 112### Data Fields113 114- `task_id`: identifier for the data sample115- `prompt`: input for the model containing function header and docstrings116- `canonical_solution`: solution for the problem in the `prompt`117- `description`: task description118- `test`: contains function to test generated code for correctness119- `entry_point`: entry point for test120- `language`: programming lanuage identifier to call the appropriate subprocess call for program execution121 122 123### Data Splits124 125 - HumanXEval126   - Python 127   - Java 128   - JavaScript129   - Csharp130   - CPP131   - Go132   - Kotlin133   - PHP134   - Perl135   - Ruby136   - Swift137   - Scala138 - MBXP139   - Python140   - Java141   - JavaScript142   - TypeScript143   - Csharp144   - CPP145   - Go146   - Kotlin147   - PHP148   - Perl149   - Ruby150   - Swift151   - Scala152 - MathQA153   - Python154   - Java155   - JavaScript156 157 158## Dataset Creation159 160### Curation Rationale161 162Since code generation models are often trained on dumps of GitHub a dataset not included in the dump was necessary to properly evaluate the model. However, since this dataset was published on GitHub it is likely to be included in future dumps.163 164### Personal and Sensitive Information165 166None.167 168### Social Impact of Dataset169With this dataset code generating models can be better evaluated which leads to fewer issues introduced when using such models.170 171### Dataset Curators172AWS AI Labs173 174## Execution175 176### Execution Example177Install the repo [mxeval](https://github.com/amazon-science/mxeval) to execute generations or canonical solutions for the prompts from this dataset.178 179```python180>>> from datasets import load_dataset181>>> from mxeval.execution import check_correctness182>>> mbxp_python = load_dataset("AmazonScience/mxeval", "mbxp", split="python")183>>> example_problem = mbxp_python[0]184>>> check_correctness(example_problem, example_problem["canonical_solution"], timeout=20.0)185{'task_id': 'MBPP/1', 'passed': True, 'result': 'passed', 'completion_id': None, 'time_elapsed': 10.582208633422852}186```187### Considerations for Using the Data188Make sure to sandbox the execution environment since generated code samples can be harmful.189 190 191### Licensing Information192 193[LICENSE](https://huggingface.co/datasets/AmazonScience/mxeval/blob/main/LICENSE) <br>194[THIRD PARTY LICENSES](https://huggingface.co/datasets/AmazonScience/mxeval/blob/main/THIRD_PARTY_LICENSES)195 196# Citation Information197```198@article{mbxp_athiwaratkun2022,199  title = {Multi-lingual Evaluation of Code Generation Models},200  author = {Athiwaratkun, Ben and201   Gouda, Sanjay Krishna and202   Wang, Zijian and203   Li, Xiaopeng and204   Tian, Yuchen and205   Tan, Ming206   and Ahmad, Wasi Uddin and207   Wang, Shiqi and208   Sun, Qing and209   Shang, Mingyue and210   Gonugondla, Sujan Kumar and211   Ding, Hantian and212   Kumar, Varun and213   Fulton, Nathan and214   Farahani, Arash and215   Jain, Siddhartha and216   Giaquinto, Robert and217   Qian, Haifeng and218   Ramanathan, Murali Krishna and219   Nallapati, Ramesh and220   Ray, Baishakhi and221   Bhatia, Parminder and222   Sengupta, Sudipta and223   Roth, Dan and224   Xiang, Bing},225  doi = {10.48550/ARXIV.2210.14868},226  url = {https://arxiv.org/abs/2210.14868},227  keywords = {Machine Learning (cs.LG), Computation and Language (cs.CL), FOS: Computer and information sciences, FOS: Computer and information sciences},228  publisher = {arXiv},229  year = {2022},230  copyright = {Creative Commons Attribution 4.0 International}231}232```233 234# Contributions235 236[skgouda@](https://github.com/sk-g) [benathi@](https://github.com/benathi)