Team Ai
Apppublic

codeparrot/code-generation-models

sourceHugging Faceapache-2.0updated 3y agoView on Hugging Face
257likes
demo_humaneval.md55 linesDownload Raw Back to evaluation
1 2We can load HumanEval dataset and pass@k metric from 🤗 [`datasets`](https://huggingface.co/docs/datasets/index) and 🤗 [`evaluate`](https://huggingface.co/docs/evaluate/index)3 4```python5from datasets import load_dataset6from evaluate import load7 8human_eval = load_dataset("openai_humaneval")9code_eval_metric = load("code_eval")10```11 12We can easily compute the pass@k for a problem that asks for the implementation of a function that sums two integers:13 14```python15test_cases = ["assert add(2,3)==5"]16candidates = [["def add(a,b): return a*b", "def add(a, b): return a+b"]]17pass_at_k, results = code_eval_metric.compute(references=test_cases, predictions=candidates, k=[1, 2])18print(pass_at_k)19{'pass@1': 0.5, 'pass@2': 1.0}20```21 22To better understand how pass@k metric works, we will illustrate it with a concrete example from HumanEval dataset. We select the problem below and see how CodeParrot 🦜 (110M) performs and which code completions pass the unit tests:23 24**Problem:**25 26```python27 28def truncate_number(number: float) -> float:29    """ Given a positive floating point number, it can be decomposed into30    and integer part (largest integer smaller than given number) and decimals31    (leftover part always smaller than 1).32 33    Return the decimal part of the number.34    >>> truncate_number(3.5)35    0.536    """37````38 39Instead of 200 candidate solutions, we will only generate 20 samples for illustration purposes. We use nucleus sampling with top-p where `p=0.95`, `temperature=0.2`, and sample tokens from the model until we encounter a stop sequence indicating the end of a method: ‘\nclass’, ‘\ndef’, ‘\n#’, ‘\nif’, or ‘\nprint’. For more details about decoding strategies for language generation, we recommend this [blog](https://huggingface.co/blog/how-to-generate).40 41**Remark**:42 43Regarding the temperature parameter, in [Codex](https://arxiv.org/pdf/2107.03374.pdf) paper, the authors observed that the best performing temperature increases as the number of samples permitted k increases. Similar behavior was also observed in [CodeGen](https://arxiv.org/pdf/2203.13474.pdf). When a model is only allowed a few samples to pass unit tests, it is beneficial to use the learned distribution, through a low temperature, to select candidates that are likely to pass. But when a model is allowed for more chances with a high k, using a higher sampling temperature to tilt the learned model distribution lets it explore diverse samples and thus have a greater chance of synthesizing a correct program.44 45 46For our experiment, we compute pass@1, pass@10 and pass@20, each corresponding to unit test pass rate when selecting respectively 1, 10 and 20 samples from the candidate solutions.47 48```49 50Results: {'pass@1': 0.1, 'pass@10': 0.7631, 'pass@20': 1.0}51 52````53 54If we take a closer look at the unit test results for each candidate solution, we find that 2 passed the unit test. This means that we have 2 correct solutions among 20, which corresponds to our pass@1 value `2/20 = 0.1`. The scores pass@10 and pass@20 are higher, because the more samples we select from the candidate completions, the more likely we are to include the correct implementation. As55for pass@20, it is `1`, since if we select all 20 candidates the problem gets solved which gives 100% success rate.