code-search-net/code_search_net
Dataset Card for CodeSearchNet corpus Dataset Summary CodeSearchNet corpus is a dataset of 2 milllion (comment, code) pairs from opensource libraries hosted on GitHub. It contains code and documentation for several programming languages. CodeSearchNet corpus was gathered to support the CodeSearchNet challenge, to explore the problem of code retrieval using natural language. Supported Tasks and Leaderboards language-modeling: The dataset can be used… See the full description on the dataset page: https://huggingface.co/datasets/code-search-net/code_search_net.
33833k
1---2annotations_creators:3- no-annotation4language_creators:5- machine-generated6language:7- code8license:9- other10multilinguality:11- multilingual12size_categories:13- 100K<n<1M14- 10K<n<100K15- 1M<n<10M16source_datasets:17- original18task_categories:19- text-generation20- fill-mask21task_ids:22- language-modeling23- masked-language-modeling24paperswithcode_id: codesearchnet25pretty_name: CodeSearchNet26dataset_info:27- config_name: all28 features:29 - name: repository_name30 dtype: string31 - name: func_path_in_repository32 dtype: string33 - name: func_name34 dtype: string35 - name: whole_func_string36 dtype: string37 - name: language38 dtype: string39 - name: func_code_string40 dtype: string41 - name: func_code_tokens42 sequence: string43 - name: func_documentation_string44 dtype: string45 - name: func_documentation_tokens46 sequence: string47 - name: split_name48 dtype: string49 - name: func_code_url50 dtype: string51 splits:52 - name: train53 num_bytes: 585059425554 num_examples: 188085355 - name: test56 num_bytes: 30862576157 num_examples: 10052958 - name: validation59 num_bytes: 27456391460 num_examples: 8915461 download_size: 196322019762 dataset_size: 643378393063- config_name: go64 features:65 - name: repository_name66 dtype: string67 - name: func_path_in_repository68 dtype: string69 - name: func_name70 dtype: string71 - name: whole_func_string72 dtype: string73 - name: language74 dtype: string75 - name: func_code_string76 dtype: string77 - name: func_code_tokens78 sequence: string79 - name: func_documentation_string80 dtype: string81 - name: func_documentation_tokens82 sequence: string83 - name: split_name84 dtype: string85 - name: func_code_url86 dtype: string87 splits:88 - name: train89 num_bytes: 73815157090 num_examples: 31783291 - name: test92 num_bytes: 3228689493 num_examples: 1429194 - name: validation95 num_bytes: 2688842396 num_examples: 1424297 download_size: 22852072698 dataset_size: 79732688799- config_name: java100 features:101 - name: repository_name102 dtype: string103 - name: func_path_in_repository104 dtype: string105 - name: func_name106 dtype: string107 - name: whole_func_string108 dtype: string109 - name: language110 dtype: string111 - name: func_code_string112 dtype: string113 - name: func_code_tokens114 sequence: string115 - name: func_documentation_string116 dtype: string117 - name: func_documentation_tokens118 sequence: string119 - name: split_name120 dtype: string121 - name: func_code_url122 dtype: string123 splits:124 - name: train125 num_bytes: 1429270143126 num_examples: 454451127 - name: test128 num_bytes: 82377090129 num_examples: 26909130 - name: validation131 num_bytes: 42358211132 num_examples: 15328133 download_size: 425659927134 dataset_size: 1554005444135- config_name: javascript136 features:137 - name: repository_name138 dtype: string139 - name: func_path_in_repository140 dtype: string141 - name: func_name142 dtype: string143 - name: whole_func_string144 dtype: string145 - name: language146 dtype: string147 - name: func_code_string148 dtype: string149 - name: func_code_tokens150 sequence: string151 - name: func_documentation_string152 dtype: string153 - name: func_documentation_tokens154 sequence: string155 - name: split_name156 dtype: string157 - name: func_code_url158 dtype: string159 splits:160 - name: train161 num_bytes: 480285847162 num_examples: 123889163 - name: test164 num_bytes: 24056920165 num_examples: 6483166 - name: validation167 num_bytes: 30168190168 num_examples: 8253169 download_size: 177637817170 dataset_size: 534510957171- config_name: php172 features:173 - name: repository_name174 dtype: string175 - name: func_path_in_repository176 dtype: string177 - name: func_name178 dtype: string179 - name: whole_func_string180 dtype: string181 - name: language182 dtype: string183 - name: func_code_string184 dtype: string185 - name: func_code_tokens186 sequence: string187 - name: func_documentation_string188 dtype: string189 - name: func_documentation_tokens190 sequence: string191 - name: split_name192 dtype: string193 - name: func_code_url194 dtype: string195 splits:196 - name: train197 num_bytes: 1532562114198 num_examples: 523712199 - name: test200 num_bytes: 80203721201 num_examples: 28391202 - name: validation203 num_bytes: 78163768204 num_examples: 26015205 download_size: 507426963206 dataset_size: 1690929603207- config_name: python208 features:209 - name: repository_name210 dtype: string211 - name: func_path_in_repository212 dtype: string213 - name: func_name214 dtype: string215 - name: whole_func_string216 dtype: string217 - name: language218 dtype: string219 - name: func_code_string220 dtype: string221 - name: func_code_tokens222 sequence: string223 - name: func_documentation_string224 dtype: string225 - name: func_documentation_tokens226 sequence: string227 - name: split_name228 dtype: string229 - name: func_code_url230 dtype: string231 splits:232 - name: train233 num_bytes: 1559643126234 num_examples: 412178235 - name: test236 num_bytes: 84341908237 num_examples: 22176238 - name: validation239 num_bytes: 92154630240 num_examples: 23107241 download_size: 581272659242 dataset_size: 1736139664243- config_name: ruby244 features:245 - name: repository_name246 dtype: string247 - name: func_path_in_repository248 dtype: string249 - name: func_name250 dtype: string251 - name: whole_func_string252 dtype: string253 - name: language254 dtype: string255 - name: func_code_string256 dtype: string257 - name: func_code_tokens258 sequence: string259 - name: func_documentation_string260 dtype: string261 - name: func_documentation_tokens262 sequence: string263 - name: split_name264 dtype: string265 - name: func_code_url266 dtype: string267 splits:268 - name: train269 num_bytes: 110681455270 num_examples: 48791271 - name: test272 num_bytes: 5359228273 num_examples: 2279274 - name: validation275 num_bytes: 4830692276 num_examples: 2209277 download_size: 42266439278 dataset_size: 120871375279config_names:280- all281- go282- java283- javascript284- php285- python286- ruby287configs:288- config_name: all289 data_files:290 - split: train291 path: all/train-*292 - split: test293 path: all/test-*294 - split: validation295 path: all/validation-*296 default: true297- config_name: go298 data_files:299 - split: train300 path: go/train-*301 - split: test302 path: go/test-*303 - split: validation304 path: go/validation-*305- config_name: java306 data_files:307 - split: train308 path: java/train-*309 - split: test310 path: java/test-*311 - split: validation312 path: java/validation-*313- config_name: javascript314 data_files:315 - split: train316 path: javascript/train-*317 - split: test318 path: javascript/test-*319 - split: validation320 path: javascript/validation-*321- config_name: php322 data_files:323 - split: train324 path: php/train-*325 - split: test326 path: php/test-*327 - split: validation328 path: php/validation-*329- config_name: python330 data_files:331 - split: train332 path: python/train-*333 - split: test334 path: python/test-*335 - split: validation336 path: python/validation-*337- config_name: ruby338 data_files:339 - split: train340 path: ruby/train-*341 - split: test342 path: ruby/test-*343 - split: validation344 path: ruby/validation-*345---346 347# Dataset Card for CodeSearchNet corpus348 349## Table of Contents350- [Dataset Description](#dataset-description)351 - [Dataset Summary](#dataset-summary)352 - [Supported Tasks and Leaderboards](#supported-tasks-and-leaderboards)353 - [Languages](#languages)354- [Dataset Structure](#dataset-structure)355 - [Data Instances](#data-instances)356 - [Data Fields](#data-fields)357 - [Data Splits](#data-splits)358- [Dataset Creation](#dataset-creation)359 - [Curation Rationale](#curation-rationale)360 - [Source Data](#source-data)361 - [Annotations](#annotations)362 - [Personal and Sensitive Information](#personal-and-sensitive-information)363- [Considerations for Using the Data](#considerations-for-using-the-data)364 - [Social Impact of Dataset](#social-impact-of-dataset)365 - [Discussion of Biases](#discussion-of-biases)366 - [Other Known Limitations](#other-known-limitations)367- [Additional Information](#additional-information)368 - [Dataset Curators](#dataset-curators)369 - [Licensing Information](#licensing-information)370 - [Citation Information](#citation-information)371 - [Contributions](#contributions)372 373## Dataset Description374- **Homepage:** https://wandb.ai/github/CodeSearchNet/benchmark375- **Repository:** https://github.com/github/CodeSearchNet376- **Paper:** https://arxiv.org/abs/1909.09436377- **Data:** https://doi.org/10.5281/zenodo.7908468378- **Leaderboard:** https://wandb.ai/github/CodeSearchNet/benchmark/leaderboard379 380### Dataset Summary381 382CodeSearchNet corpus is a dataset of 2 milllion (comment, code) pairs from opensource libraries hosted on GitHub. It contains code and documentation for several programming languages.383 384CodeSearchNet corpus was gathered to support the [CodeSearchNet challenge](https://wandb.ai/github/CodeSearchNet/benchmark), to explore the problem of code retrieval using natural language.385 386### Supported Tasks and Leaderboards387 388- `language-modeling`: The dataset can be used to train a model for modelling programming languages, which consists in building language models for programming languages.389 390### Languages391 392- Go **programming** language393- Java **programming** language394- Javascript **programming** language395- PHP **programming** language396- Python **programming** language397- Ruby **programming** language398 399## Dataset Structure400 401### Data Instances402 403A data point consists of a function code along with its documentation. Each data point also contains meta data on the function, such as the repository it was extracted from.404```405{406 'id': '0',407 'repository_name': 'organisation/repository',408 'func_path_in_repository': 'src/path/to/file.py',409 'func_name': 'func',410 'whole_func_string': 'def func(args):\n"""Docstring"""\n [...]',411 'language': 'python', 412 'func_code_string': '[...]',413 'func_code_tokens': ['def', 'func', '(', 'args', ')', ...],414 'func_documentation_string': 'Docstring',415 'func_documentation_string_tokens': ['Docstring'],416 'split_name': 'train',417 'func_code_url': 'https://github.com/<org>/<repo>/blob/<hash>/src/path/to/file.py#L111-L150'418}419```420### Data Fields421 422- `id`: Arbitrary number423- `repository_name`: name of the GitHub repository424- `func_path_in_repository`: tl;dr: path to the file which holds the function in the repository425- `func_name`: name of the function in the file426- `whole_func_string`: Code + documentation of the function427- `language`: Programming language in whoch the function is written428- `func_code_string`: Function code429- `func_code_tokens`: Tokens yielded by Treesitter430- `func_documentation_string`: Function documentation431- `func_documentation_string_tokens`: Tokens yielded by Treesitter432- `split_name`: Name of the split to which the example belongs (one of train, test or valid)433- `func_code_url`: URL to the function code on Github434 435### Data Splits436 437Three splits are available:438- train439- test440- valid441 442## Dataset Creation443 444### Curation Rationale445 446[More Information Needed]447 448### Source Data449 450#### Initial Data Collection and Normalization451 452All information can be retrieved in the [original technical review](https://arxiv.org/pdf/1909.09436.pdf)453 454**Corpus collection**:455 456Corpus has been collected from publicly available open-source non-fork GitHub repositories, using libraries.io to identify all projects which are used by at least one other project, and sort them by “popularity” as indicated by the number of stars and forks. 457 458Then, any projects that do not have a license or whose license does not explicitly permit the re-distribution of parts of the project were removed. Treesitter - GitHub's universal parser - has been used to then tokenize all Go, Java, JavaScript, Python, PHP and Ruby functions (or methods) using and, where available, their respective documentation text using a heuristic regular expression.459 460**Corpus filtering**:461 462Functions without documentation are removed from the corpus. This yields a set of pairs ($c_i$, $d_i$) where ci is some function documented by di. Pairs ($c_i$, $d_i$) are passed through the folllowing preprocessing tasks:463 464- Documentation $d_i$ is truncated to the first full paragraph to remove in-depth discussion of function arguments and return values465- Pairs in which $d_i$ is shorter than three tokens are removed466- Functions $c_i$ whose implementation is shorter than three lines are removed467- Functions whose name contains the substring “test” are removed468- Constructors and standard extenion methods (eg `__str__` in Python or `toString` in Java) are removed469- Duplicates and near duplicates functions are removed, in order to keep only one version of the function470 471#### Who are the source language producers?472 473OpenSource contributors produced the code and documentations.474 475The dataset was gatherered and preprocessed automatically.476 477### Annotations478 479#### Annotation process480 481[More Information Needed]482 483#### Who are the annotators?484 485[More Information Needed]486 487### Personal and Sensitive Information488 489[More Information Needed]490 491## Considerations for Using the Data492 493### Social Impact of Dataset494 495[More Information Needed]496 497### Discussion of Biases498 499[More Information Needed]500 501### Other Known Limitations502 503[More Information Needed]504 505## Additional Information506 507### Dataset Curators508 509[More Information Needed]510 511### Licensing Information512 513Each example in the dataset has is extracted from a GitHub repository, and each repository has its own license. Example-wise license information is not (yet) included in this dataset: you will need to find out yourself which license the code is using.514 515### Citation Information516 517@article{husain2019codesearchnet,518 title={{CodeSearchNet} challenge: Evaluating the state of semantic code search},519 author={Husain, Hamel and Wu, Ho-Hsiang and Gazit, Tiferet and Allamanis, Miltiadis and Brockschmidt, Marc},520 journal={arXiv preprint arXiv:1909.09436},521 year={2019}522}523 524### Contributions525 526Thanks to [@SBrandeis](https://github.com/SBrandeis) for adding this dataset.527 