tianyang/repobench_python_v1.1
RepoBench v1.1 (Python) Introduction This dataset presents the Python portion of RepoBench v1.1 (ICLR 2024). The data encompasses a collection from GitHub, spanning the period from October 6th to December 31st, 2023. With a commitment to data integrity, we've implemented a deduplication process based on file content against the Stack v2 dataset (coming soon), aiming to mitigate data leakage and memorization concerns. Resources and Links Paper… See the full description on the dataset page: https://huggingface.co/datasets/tianyang/repobench_python_v1.1.
111.3k
1---2configs:3- config_name: default4 data_files:5 - split: cross_file_first6 path: data/cross_file_first-*7 - split: cross_file_random8 path: data/cross_file_random-*9 - split: in_file10 path: data/in_file-*11dataset_info:12 features:13 - name: repo_name14 dtype: string15 - name: file_path16 dtype: string17 - name: context18 list:19 - name: identifier20 dtype: string21 - name: path22 dtype: string23 - name: snippet24 dtype: string25 - name: import_statement26 dtype: string27 - name: token_num28 dtype: int6429 - name: cropped_code30 dtype: string31 - name: all_code32 dtype: string33 - name: next_line34 dtype: string35 - name: gold_snippet_index36 dtype: int6437 - name: created_at38 dtype: string39 - name: level40 dtype: string41 splits:42 - name: cross_file_first43 num_bytes: 50452843144 num_examples: 803345 - name: cross_file_random46 num_bytes: 46724245547 num_examples: 761848 - name: in_file49 num_bytes: 48899910050 num_examples: 791051 download_size: 47299429952 dataset_size: 146076998653license: cc54task_categories:55- text-generation56language:57- en58tags:59- code60---61# RepoBench v1.1 (Python)62 63## Introduction64 65This dataset presents the **Python** portion of [RepoBench](https://arxiv.org/abs/2306.03091) v1.1 (ICLR 2024). The data encompasses a collection from GitHub, spanning the period from **October 6th to December 31st, 2023**. With a commitment to data integrity, we've implemented a deduplication process based on file content against the Stack v2 dataset (coming soon), aiming to mitigate data leakage and memorization concerns.66 67## Resources and Links68 69- [Paper](https://arxiv.org/abs/2306.03091)70- [GitHub](https://github.com/Leolty/repobench)71- [Dataset Introduction](https://github.com/Leolty/repobench/blob/main/data/README.md)72 73## FAQs74 75- **Q:** What do the features in the dataset mean?76 77 **A:** Imagine you're coding in Python and you want to write the next line of your code. The dataset provides you the following information:78 - `repo_name` (string): the name of the repository79 - `file_path` (string): the path of the current file80 - `context` (list): the cross-file code snippets that might be helpful for writing the next line:81 - `identifier` (string): the identifier of the code snippet82 - `path` (string): the path of the code snippet83 - `snippet` (string): the code snippet84 - `import_statement` (string): the import statement of the current file85 - `cropped_code` (string): the cropped code of the current file (up to previous 120 lines)86 - `all_code` (string): the entire code of the current file (not cropped)87 - `next_line` (string): the next line of the code (this serves as the target)88 - `gold_snippet_index` (int): the index of the gold snippet in the context (which will be used in next line, just for reference, you should not use this for next line prediction)89 - `created_at` (string): the creation time of the repository90 - `level` (string): the level of next line completion, which is measured by the number of tokens for the whole prompt (including all the context, import statement, cropped code and some neccessary separator tokens)91 92- **Q:** How does the level be defined?93 94 **A:** The level is determined by the number of tokens for the whole prompt (including all the context, import statement, cropped code and some neccessary separator tokens). The token number is calculated by the tokenizer of GPT-4 by using [tiktoken](https://github.com/openai/tiktoken). The following table shows the level definition:95 96 | Level | Prompt Length (Number of Tokens) |97 |-------|------------------------|98 | 2k | 640 - 1,600 |99 | 4k | 1,600 - 3,600 |100 | 8k | 3,600 - 7,200 |101 | 12k | 7,200 - 10,800 |102 | 16k | 10,800 - 14,400 |103 | 24k | 14,400 - 21,600 |104 | 32k | 21,600 - 28,800 |105 | 64k | 28,800 - 57,600 |106 | 128k | 57,600 - 100,000 |107 108- **Q:** What does the different splits mean?109 110 **A:** The dataset is split into three parts:111 - `cross_file_first`: the next line of code utilizes content from a cross-file code snippet and it is its first usage within current file.112 - `cross_file_random`: the next line of code utilizes content from a cross-file code snippet and it is NOT its first usage within current file.113 - `in_file`: the next line of code does not utilize content from a cross-file code snippet.114 115- **Q:** How to construct the prompt for next line prediction?116 117 **A:** We hereby provide the official implementation for constructing prompts. Please note that the methods described below are not necessarily the optimal way of construction. Reordering, retrieval argumentation, or employing different cropping/construction techniques could potentially lead to varying degrees of improvement. Ensure that your model evaluations are conducted in a fair manner.118 119 ```python120 import re121 122 def construct_prompt(123 data: dict, 124 language: str = "python",125 tokenizer= None,126 max_token_nums: int = 15800127 ) -> str:128 """129 Construct the prompt for next line prediction.130 131 :param data: data point from the dataset132 :param language: the language of the code133 :param tokenizer: the tokenizer of the evaluation model134 :param max_token_nums: the maximum number of tokens constraint for the prompt135 136 :return: the constructed prompt137 """138 139 # comment symbol for different languages140 comment_symbol = "#" if language == "python" else "//"141 142 # construct the cross-file prompt and in-file prompt separately143 # cross-file prompt144 cross_file_prompt = f"{comment_symbol} Repo Name: {data['repo_name']}\n"145 146 for snippet in data['context']:147 cross_file_prompt += f"{comment_symbol} Path: {snippet['path']}\n{snippet['snippet']}" + "\n\n"148 149 # in-file prompt150 in_file_prompt = f"{comment_symbol} Path: {data['file_path']}\n{data['import_statement']}\n{data['cropped_code']}\n"151 152 # if we assign the tokenizer and the max_token_nums, we will truncate the cross-file prompt to meet the constraint153 if tokenizer is not None and max_token_nums is not None:154 155 cross_file_prompt_token_nums = len(tokenizer.encode(cross_file_prompt))156 in_file_prompt_token_nums = len(tokenizer.encode(in_file_prompt))157 158 exceed_token_nums = cross_file_prompt_token_nums + in_file_prompt_token_nums - max_token_nums159 160 if exceed_token_nums > 0:161 # split the cross-file prompt into lines162 cross_file_prompt_lines = cross_file_prompt.split("\n")163 # drop lines from end until the extra token number is less than 0164 for i in range(len(repo_prompt_lines)-1, -1, -1):165 extra_token_num -= len(tokenizer.encode(cross_file_prompt_lines[i]))166 if extra_token_num < 0:167 break168 169 # join the lines back170 cross_file_prompt = "\n".join(cross_file_prompt_lines[:i]) + "\n\n"171 172 # combine the cross-file prompt and in-file prompt173 prompt = cross_file_prompt + in_file_prompt174 175 # normalize some empty lines176 prompt = re.sub(r'\n{4,}', '\n\n', prompt)177 178 return prompt179 ```180 181- **Q:** How to load the dataset?182 183 **A:** You can simply use the following code to load the dataset:184 185 ```python186 from datasets import load_dataset187 188 dataset = load_dataset("tianyang/repobench_python_v1.1")189 ```190 191 To construct the prompt for next line prediction, you can refer to the official implementation provided in the previous question and use the `construct_prompt` function to construct the prompt, for example:192 193 ```python194 from transformers import AutoTokenizer, AutoModelForCausalLM195 196 tokenizer = AutoTokenizer.from_pretrained("deepseek-ai/deepseek-coder-1.3b-base")197 model = AutoModelForCausalLM.from_pretrained("deepseek-ai/deepseek-coder-1.3b-base")198 199 prompt = construct_prompt(dataset['cross_file_first'][0], tokenizer=tokenizer, max_token_nums=15800)200 ```201 202- **Q:** How often will the dataset be updated?203 204 **A:** We plan to update the dataset every three months, but there might be slight delays considering the time required for data crawling and our own schedules. If you require updated data, please feel free to contact us, and we can coordinate the timing and expedite the process.205 206- **Q:** What models should I use to evaluate the dataset?207 208 **A:** RepoBench is designed to evaluate base models, not those that have been instruction fine-tuned. Please use base models for evaluation.209 210- **Q:** I am training a new model but the knowledge cutoff date is after the dataset's. Can you provide me with the latest data?211 212 **A:** Sure! We are happy to provide you with the latest data (even customized data with specific requirements). Please feel free to contact us.213 214- **Q:** Can I opt-out?215 216 **A:** Yes, you can opt-out your repository from the dataset. Please check [Am I in RepoBench?](https://huggingface.co/spaces/tianyang/in-the-repobench), we will upload the raw data of the repository information we crawled at least 15 days before the dataset creation and release. We will respect your decision and remove your repository from the dataset if you opt-out.217 218## Citation219 220If you find RepoBench useful in your research, please consider citing the paper using the following BibTeX entry:221 222```bibtex223@misc{liu2023repobench,224 title={RepoBench: Benchmarking Repository-Level Code Auto-Completion Systems}, 225 author={Tianyang Liu and Canwen Xu and Julian McAuley},226 year={2024},227 url={https://arxiv.org/abs/2306.03091},228 booktitle={International Conference on Learning Representations}229}230```231 232Your interest and contributions to RepoBench are immensely valued. Happy coding! 🚀