Team Ai
Datasetpublic

tianyang/repobench_python_v1.1

RepoBench v1.1 (Python) Introduction This dataset presents the Python portion of RepoBench v1.1 (ICLR 2024). The data encompasses a collection from GitHub, spanning the period from October 6th to December 31st, 2023. With a commitment to data integrity, we've implemented a deduplication process based on file content against the Stack v2 dataset (coming soon), aiming to mitigate data leakage and memorization concerns. Resources and Links Paper… See the full description on the dataset page: https://huggingface.co/datasets/tianyang/repobench_python_v1.1.

sourceHugging Faceccupdated 3y agoView on Hugging Face
11likes1.3kdownloads
README.md232 linesDownload Raw Back to root
1---2configs:3- config_name: default4  data_files:5  - split: cross_file_first6    path: data/cross_file_first-*7  - split: cross_file_random8    path: data/cross_file_random-*9  - split: in_file10    path: data/in_file-*11dataset_info:12  features:13  - name: repo_name14    dtype: string15  - name: file_path16    dtype: string17  - name: context18    list:19    - name: identifier20      dtype: string21    - name: path22      dtype: string23    - name: snippet24      dtype: string25  - name: import_statement26    dtype: string27  - name: token_num28    dtype: int6429  - name: cropped_code30    dtype: string31  - name: all_code32    dtype: string33  - name: next_line34    dtype: string35  - name: gold_snippet_index36    dtype: int6437  - name: created_at38    dtype: string39  - name: level40    dtype: string41  splits:42  - name: cross_file_first43    num_bytes: 50452843144    num_examples: 803345  - name: cross_file_random46    num_bytes: 46724245547    num_examples: 761848  - name: in_file49    num_bytes: 48899910050    num_examples: 791051  download_size: 47299429952  dataset_size: 146076998653license: cc54task_categories:55- text-generation56language:57- en58tags:59- code60---61# RepoBench v1.1 (Python)62 63## Introduction64 65This dataset presents the **Python** portion of [RepoBench](https://arxiv.org/abs/2306.03091) v1.1 (ICLR 2024). The data encompasses a collection from GitHub, spanning the period from **October 6th to December 31st, 2023**. With a commitment to data integrity, we've implemented a deduplication process based on file content against the Stack v2 dataset (coming soon), aiming to mitigate data leakage and memorization concerns.66 67## Resources and Links68 69- [Paper](https://arxiv.org/abs/2306.03091)70- [GitHub](https://github.com/Leolty/repobench)71- [Dataset Introduction](https://github.com/Leolty/repobench/blob/main/data/README.md)72 73## FAQs74 75- **Q:** What do the features in the dataset mean?76  77  **A:** Imagine you're coding in Python and you want to write the next line of your code. The dataset provides you the following information:78    - `repo_name` (string): the name of the repository79    - `file_path` (string): the path of the current file80    - `context` (list): the cross-file code snippets that might be helpful for writing the next line:81      - `identifier` (string): the identifier of the code snippet82      - `path` (string): the path of the code snippet83      - `snippet` (string): the code snippet84    - `import_statement` (string): the import statement of the current file85    - `cropped_code` (string): the cropped code of the current file (up to previous 120 lines)86    - `all_code` (string): the entire code of the current file (not cropped)87    - `next_line` (string): the next line of the code (this serves as the target)88    - `gold_snippet_index` (int): the index of the gold snippet in the context (which will be used in next line, just for reference, you should not use this for next line prediction)89    - `created_at` (string): the creation time of the repository90    - `level` (string): the level of next line completion, which is measured by the number of tokens for the whole prompt (including all the context, import statement, cropped code and some neccessary separator tokens)91 92- **Q:** How does the level be defined?93 94  **A:** The level is determined by the number of tokens for the whole prompt (including all the context, import statement, cropped code and some neccessary separator tokens). The token number is calculated by the tokenizer of GPT-4 by using [tiktoken](https://github.com/openai/tiktoken). The following table shows the level definition:95 96    | Level | Prompt Length (Number of Tokens) |97    |-------|------------------------|98    | 2k    | 640 - 1,600            |99    | 4k    | 1,600 - 3,600          |100    | 8k    | 3,600 - 7,200          |101    | 12k   | 7,200 - 10,800         |102    | 16k   | 10,800 - 14,400        |103    | 24k   | 14,400 - 21,600        |104    | 32k   | 21,600 - 28,800        |105    | 64k   | 28,800 - 57,600        |106    | 128k  | 57,600 - 100,000       |107 108- **Q:** What does the different splits mean?109 110  **A:** The dataset is split into three parts:111    - `cross_file_first`: the next line of code utilizes content from a cross-file code snippet and it is its first usage within current file.112    - `cross_file_random`: the next line of code utilizes content from a cross-file code snippet and it is NOT its first usage within current file.113    - `in_file`: the next line of code does not utilize content from a cross-file code snippet.114 115- **Q:** How to construct the prompt for next line prediction?116 117  **A:** We hereby provide the official implementation for constructing prompts. Please note that the methods described below are not necessarily the optimal way of construction. Reordering, retrieval argumentation, or employing different cropping/construction techniques could potentially lead to varying degrees of improvement. Ensure that your model evaluations are conducted in a fair manner.118 119    ```python120    import re121 122    def construct_prompt(123        data: dict, 124        language: str = "python",125        tokenizer= None,126        max_token_nums: int = 15800127        ) -> str:128        """129        Construct the prompt for next line prediction.130 131        :param data: data point from the dataset132        :param language: the language of the code133        :param tokenizer: the tokenizer of the evaluation model134        :param max_token_nums: the maximum number of tokens constraint for the prompt135 136        :return: the constructed prompt137        """138 139        # comment symbol for different languages140        comment_symbol = "#" if language == "python" else "//"141 142        # construct the cross-file prompt and in-file prompt separately143        # cross-file prompt144        cross_file_prompt = f"{comment_symbol} Repo Name: {data['repo_name']}\n"145 146        for snippet in data['context']:147            cross_file_prompt += f"{comment_symbol} Path: {snippet['path']}\n{snippet['snippet']}" + "\n\n"148        149        # in-file prompt150        in_file_prompt = f"{comment_symbol} Path: {data['file_path']}\n{data['import_statement']}\n{data['cropped_code']}\n"151 152        # if we assign the tokenizer and the max_token_nums, we will truncate the cross-file prompt to meet the constraint153        if tokenizer is not None and max_token_nums is not None:154            155            cross_file_prompt_token_nums = len(tokenizer.encode(cross_file_prompt))156            in_file_prompt_token_nums = len(tokenizer.encode(in_file_prompt))157 158            exceed_token_nums = cross_file_prompt_token_nums + in_file_prompt_token_nums - max_token_nums159 160            if exceed_token_nums > 0:161                # split the cross-file prompt into lines162                cross_file_prompt_lines = cross_file_prompt.split("\n")163                # drop lines from end until the extra token number is less than 0164                for i in range(len(repo_prompt_lines)-1, -1, -1):165                    extra_token_num -= len(tokenizer.encode(cross_file_prompt_lines[i]))166                    if extra_token_num < 0:167                        break168                169                # join the lines back170                cross_file_prompt = "\n".join(cross_file_prompt_lines[:i]) + "\n\n"171        172        # combine the cross-file prompt and in-file prompt173        prompt = cross_file_prompt + in_file_prompt174 175        # normalize some empty lines176        prompt = re.sub(r'\n{4,}', '\n\n', prompt)177 178        return prompt179    ```180 181- **Q:** How to load the dataset?182 183  **A:** You can simply use the following code to load the dataset:184 185    ```python186    from datasets import load_dataset187 188    dataset = load_dataset("tianyang/repobench_python_v1.1")189    ```190 191    To construct the prompt for next line prediction, you can refer to the official implementation provided in the previous question and use the `construct_prompt` function to construct the prompt, for example:192 193    ```python194    from transformers import AutoTokenizer, AutoModelForCausalLM195 196    tokenizer = AutoTokenizer.from_pretrained("deepseek-ai/deepseek-coder-1.3b-base")197    model = AutoModelForCausalLM.from_pretrained("deepseek-ai/deepseek-coder-1.3b-base")198 199    prompt = construct_prompt(dataset['cross_file_first'][0], tokenizer=tokenizer, max_token_nums=15800)200    ```201 202- **Q:** How often will the dataset be updated?203 204  **A:** We plan to update the dataset every three months, but there might be slight delays considering the time required for data crawling and our own schedules. If you require updated data, please feel free to contact us, and we can coordinate the timing and expedite the process.205 206- **Q:** What models should I use to evaluate the dataset?207 208  **A:** RepoBench is designed to evaluate base models, not those that have been instruction fine-tuned. Please use base models for evaluation.209 210- **Q:** I am training a new model but the knowledge cutoff date is after the dataset's. Can you provide me with the latest data?211 212  **A:** Sure! We are happy to provide you with the latest data (even customized data with specific requirements). Please feel free to contact us.213 214- **Q:** Can I opt-out?215    216  **A:** Yes, you can opt-out your repository from the dataset. Please check [Am I in RepoBench?](https://huggingface.co/spaces/tianyang/in-the-repobench), we will upload the raw data of the repository information we crawled at least 15 days before the dataset creation and release. We will respect your decision and remove your repository from the dataset if you opt-out.217 218## Citation219 220If you find RepoBench useful in your research, please consider citing the paper using the following BibTeX entry:221 222```bibtex223@misc{liu2023repobench,224      title={RepoBench: Benchmarking Repository-Level Code Auto-Completion Systems}, 225      author={Tianyang Liu and Canwen Xu and Julian McAuley},226      year={2024},227      url={https://arxiv.org/abs/2306.03091},228      booktitle={International Conference on Learning Representations}229}230```231 232Your interest and contributions to RepoBench are immensely valued. Happy coding! 🚀