Team Ai
Datasetpublic

code-search-net/code_search_net

Dataset Card for CodeSearchNet corpus Dataset Summary CodeSearchNet corpus is a dataset of 2 milllion (comment, code) pairs from opensource libraries hosted on GitHub. It contains code and documentation for several programming languages. CodeSearchNet corpus was gathered to support the CodeSearchNet challenge, to explore the problem of code retrieval using natural language. Supported Tasks and Leaderboards language-modeling: The dataset can be used… See the full description on the dataset page: https://huggingface.co/datasets/code-search-net/code_search_net.

sourceHugging Faceotherupdated 8mo agoView on Hugging Face
338likes33kdownloads
README.md527 linesDownload Raw Back to root
1---2annotations_creators:3- no-annotation4language_creators:5- machine-generated6language:7- code8license:9- other10multilinguality:11- multilingual12size_categories:13- 100K<n<1M14- 10K<n<100K15- 1M<n<10M16source_datasets:17- original18task_categories:19- text-generation20- fill-mask21task_ids:22- language-modeling23- masked-language-modeling24paperswithcode_id: codesearchnet25pretty_name: CodeSearchNet26dataset_info:27- config_name: all28  features:29  - name: repository_name30    dtype: string31  - name: func_path_in_repository32    dtype: string33  - name: func_name34    dtype: string35  - name: whole_func_string36    dtype: string37  - name: language38    dtype: string39  - name: func_code_string40    dtype: string41  - name: func_code_tokens42    sequence: string43  - name: func_documentation_string44    dtype: string45  - name: func_documentation_tokens46    sequence: string47  - name: split_name48    dtype: string49  - name: func_code_url50    dtype: string51  splits:52  - name: train53    num_bytes: 585059425554    num_examples: 188085355  - name: test56    num_bytes: 30862576157    num_examples: 10052958  - name: validation59    num_bytes: 27456391460    num_examples: 8915461  download_size: 196322019762  dataset_size: 643378393063- config_name: go64  features:65  - name: repository_name66    dtype: string67  - name: func_path_in_repository68    dtype: string69  - name: func_name70    dtype: string71  - name: whole_func_string72    dtype: string73  - name: language74    dtype: string75  - name: func_code_string76    dtype: string77  - name: func_code_tokens78    sequence: string79  - name: func_documentation_string80    dtype: string81  - name: func_documentation_tokens82    sequence: string83  - name: split_name84    dtype: string85  - name: func_code_url86    dtype: string87  splits:88  - name: train89    num_bytes: 73815157090    num_examples: 31783291  - name: test92    num_bytes: 3228689493    num_examples: 1429194  - name: validation95    num_bytes: 2688842396    num_examples: 1424297  download_size: 22852072698  dataset_size: 79732688799- config_name: java100  features:101  - name: repository_name102    dtype: string103  - name: func_path_in_repository104    dtype: string105  - name: func_name106    dtype: string107  - name: whole_func_string108    dtype: string109  - name: language110    dtype: string111  - name: func_code_string112    dtype: string113  - name: func_code_tokens114    sequence: string115  - name: func_documentation_string116    dtype: string117  - name: func_documentation_tokens118    sequence: string119  - name: split_name120    dtype: string121  - name: func_code_url122    dtype: string123  splits:124  - name: train125    num_bytes: 1429270143126    num_examples: 454451127  - name: test128    num_bytes: 82377090129    num_examples: 26909130  - name: validation131    num_bytes: 42358211132    num_examples: 15328133  download_size: 425659927134  dataset_size: 1554005444135- config_name: javascript136  features:137  - name: repository_name138    dtype: string139  - name: func_path_in_repository140    dtype: string141  - name: func_name142    dtype: string143  - name: whole_func_string144    dtype: string145  - name: language146    dtype: string147  - name: func_code_string148    dtype: string149  - name: func_code_tokens150    sequence: string151  - name: func_documentation_string152    dtype: string153  - name: func_documentation_tokens154    sequence: string155  - name: split_name156    dtype: string157  - name: func_code_url158    dtype: string159  splits:160  - name: train161    num_bytes: 480285847162    num_examples: 123889163  - name: test164    num_bytes: 24056920165    num_examples: 6483166  - name: validation167    num_bytes: 30168190168    num_examples: 8253169  download_size: 177637817170  dataset_size: 534510957171- config_name: php172  features:173  - name: repository_name174    dtype: string175  - name: func_path_in_repository176    dtype: string177  - name: func_name178    dtype: string179  - name: whole_func_string180    dtype: string181  - name: language182    dtype: string183  - name: func_code_string184    dtype: string185  - name: func_code_tokens186    sequence: string187  - name: func_documentation_string188    dtype: string189  - name: func_documentation_tokens190    sequence: string191  - name: split_name192    dtype: string193  - name: func_code_url194    dtype: string195  splits:196  - name: train197    num_bytes: 1532562114198    num_examples: 523712199  - name: test200    num_bytes: 80203721201    num_examples: 28391202  - name: validation203    num_bytes: 78163768204    num_examples: 26015205  download_size: 507426963206  dataset_size: 1690929603207- config_name: python208  features:209  - name: repository_name210    dtype: string211  - name: func_path_in_repository212    dtype: string213  - name: func_name214    dtype: string215  - name: whole_func_string216    dtype: string217  - name: language218    dtype: string219  - name: func_code_string220    dtype: string221  - name: func_code_tokens222    sequence: string223  - name: func_documentation_string224    dtype: string225  - name: func_documentation_tokens226    sequence: string227  - name: split_name228    dtype: string229  - name: func_code_url230    dtype: string231  splits:232  - name: train233    num_bytes: 1559643126234    num_examples: 412178235  - name: test236    num_bytes: 84341908237    num_examples: 22176238  - name: validation239    num_bytes: 92154630240    num_examples: 23107241  download_size: 581272659242  dataset_size: 1736139664243- config_name: ruby244  features:245  - name: repository_name246    dtype: string247  - name: func_path_in_repository248    dtype: string249  - name: func_name250    dtype: string251  - name: whole_func_string252    dtype: string253  - name: language254    dtype: string255  - name: func_code_string256    dtype: string257  - name: func_code_tokens258    sequence: string259  - name: func_documentation_string260    dtype: string261  - name: func_documentation_tokens262    sequence: string263  - name: split_name264    dtype: string265  - name: func_code_url266    dtype: string267  splits:268  - name: train269    num_bytes: 110681455270    num_examples: 48791271  - name: test272    num_bytes: 5359228273    num_examples: 2279274  - name: validation275    num_bytes: 4830692276    num_examples: 2209277  download_size: 42266439278  dataset_size: 120871375279config_names:280- all281- go282- java283- javascript284- php285- python286- ruby287configs:288- config_name: all289  data_files:290  - split: train291    path: all/train-*292  - split: test293    path: all/test-*294  - split: validation295    path: all/validation-*296  default: true297- config_name: go298  data_files:299  - split: train300    path: go/train-*301  - split: test302    path: go/test-*303  - split: validation304    path: go/validation-*305- config_name: java306  data_files:307  - split: train308    path: java/train-*309  - split: test310    path: java/test-*311  - split: validation312    path: java/validation-*313- config_name: javascript314  data_files:315  - split: train316    path: javascript/train-*317  - split: test318    path: javascript/test-*319  - split: validation320    path: javascript/validation-*321- config_name: php322  data_files:323  - split: train324    path: php/train-*325  - split: test326    path: php/test-*327  - split: validation328    path: php/validation-*329- config_name: python330  data_files:331  - split: train332    path: python/train-*333  - split: test334    path: python/test-*335  - split: validation336    path: python/validation-*337- config_name: ruby338  data_files:339  - split: train340    path: ruby/train-*341  - split: test342    path: ruby/test-*343  - split: validation344    path: ruby/validation-*345---346 347# Dataset Card for CodeSearchNet corpus348 349## Table of Contents350- [Dataset Description](#dataset-description)351  - [Dataset Summary](#dataset-summary)352  - [Supported Tasks and Leaderboards](#supported-tasks-and-leaderboards)353  - [Languages](#languages)354- [Dataset Structure](#dataset-structure)355  - [Data Instances](#data-instances)356  - [Data Fields](#data-fields)357  - [Data Splits](#data-splits)358- [Dataset Creation](#dataset-creation)359  - [Curation Rationale](#curation-rationale)360  - [Source Data](#source-data)361  - [Annotations](#annotations)362  - [Personal and Sensitive Information](#personal-and-sensitive-information)363- [Considerations for Using the Data](#considerations-for-using-the-data)364  - [Social Impact of Dataset](#social-impact-of-dataset)365  - [Discussion of Biases](#discussion-of-biases)366  - [Other Known Limitations](#other-known-limitations)367- [Additional Information](#additional-information)368  - [Dataset Curators](#dataset-curators)369  - [Licensing Information](#licensing-information)370  - [Citation Information](#citation-information)371  - [Contributions](#contributions)372 373## Dataset Description374- **Homepage:** https://wandb.ai/github/CodeSearchNet/benchmark375- **Repository:** https://github.com/github/CodeSearchNet376- **Paper:** https://arxiv.org/abs/1909.09436377- **Data:** https://doi.org/10.5281/zenodo.7908468378- **Leaderboard:** https://wandb.ai/github/CodeSearchNet/benchmark/leaderboard379 380### Dataset Summary381 382CodeSearchNet corpus is a dataset of 2 milllion (comment, code) pairs from opensource libraries hosted on GitHub. It contains code and documentation for several programming languages.383 384CodeSearchNet corpus was gathered to support the [CodeSearchNet challenge](https://wandb.ai/github/CodeSearchNet/benchmark), to explore the problem of code retrieval using natural language.385 386### Supported Tasks and Leaderboards387 388- `language-modeling`: The dataset can be used to train a model for modelling programming languages, which consists in building language models for programming languages.389 390### Languages391 392- Go **programming** language393- Java **programming** language394- Javascript **programming** language395- PHP **programming** language396- Python **programming** language397- Ruby **programming** language398 399## Dataset Structure400 401### Data Instances402 403A data point consists of a function code along with its documentation. Each data point also contains meta data on the function, such as the repository it was extracted from.404```405{406  'id': '0',407  'repository_name': 'organisation/repository',408  'func_path_in_repository': 'src/path/to/file.py',409  'func_name': 'func',410  'whole_func_string': 'def func(args):\n"""Docstring"""\n [...]',411  'language': 'python', 412  'func_code_string': '[...]',413  'func_code_tokens': ['def', 'func', '(', 'args', ')', ...],414  'func_documentation_string': 'Docstring',415  'func_documentation_string_tokens': ['Docstring'],416  'split_name': 'train',417  'func_code_url': 'https://github.com/<org>/<repo>/blob/<hash>/src/path/to/file.py#L111-L150'418}419```420### Data Fields421 422- `id`: Arbitrary number423- `repository_name`: name of the GitHub repository424- `func_path_in_repository`: tl;dr: path to the file which holds the function in the repository425- `func_name`: name of the function in the file426- `whole_func_string`: Code + documentation of the function427- `language`: Programming language in whoch the function is written428- `func_code_string`: Function code429- `func_code_tokens`: Tokens yielded by Treesitter430- `func_documentation_string`: Function documentation431- `func_documentation_string_tokens`: Tokens yielded by Treesitter432- `split_name`: Name of the split to which the example belongs (one of train, test or valid)433- `func_code_url`: URL to the function code on Github434 435### Data Splits436 437Three splits are available:438- train439- test440- valid441 442## Dataset Creation443 444### Curation Rationale445 446[More Information Needed]447 448### Source Data449 450#### Initial Data Collection and Normalization451 452All information can be retrieved in the [original technical review](https://arxiv.org/pdf/1909.09436.pdf)453 454**Corpus collection**:455 456Corpus has been collected from publicly available open-source non-fork GitHub repositories, using libraries.io to identify all projects which are used by at least one other project, and sort them by “popularity” as indicated by the number of stars and forks. 457 458Then, any projects that do not have a license or whose license does not explicitly permit the re-distribution of parts of the project were removed. Treesitter - GitHub's universal parser - has been used to then tokenize all Go, Java, JavaScript, Python, PHP and Ruby functions (or methods) using and, where available, their respective documentation text using a heuristic regular expression.459 460**Corpus filtering**:461 462Functions without documentation are removed from the corpus. This yields a set of pairs ($c_i$, $d_i$) where ci is some function documented by di. Pairs ($c_i$, $d_i$) are passed through the folllowing preprocessing tasks:463 464- Documentation $d_i$ is truncated to the first full paragraph to remove in-depth discussion of function arguments and return values465- Pairs in which $d_i$ is shorter than three tokens are removed466- Functions $c_i$ whose implementation is shorter than three lines are removed467- Functions whose name contains the substring “test” are removed468- Constructors and standard extenion methods (eg `__str__` in Python or `toString` in Java) are removed469- Duplicates and near duplicates functions are removed, in order to keep only one version of the function470 471#### Who are the source language producers?472 473OpenSource contributors produced the code and documentations.474 475The dataset was gatherered and preprocessed automatically.476 477### Annotations478 479#### Annotation process480 481[More Information Needed]482 483#### Who are the annotators?484 485[More Information Needed]486 487### Personal and Sensitive Information488 489[More Information Needed]490 491## Considerations for Using the Data492 493### Social Impact of Dataset494 495[More Information Needed]496 497### Discussion of Biases498 499[More Information Needed]500 501### Other Known Limitations502 503[More Information Needed]504 505## Additional Information506 507### Dataset Curators508 509[More Information Needed]510 511### Licensing Information512 513Each example in the dataset has is extracted from a GitHub repository, and each repository has its own license. Example-wise license information is not (yet) included in this dataset: you will need to find out yourself which license the code is using.514 515### Citation Information516 517@article{husain2019codesearchnet,518  title={{CodeSearchNet} challenge: Evaluating the state of semantic code search},519  author={Husain, Hamel and Wu, Ho-Hsiang and Gazit, Tiferet and Allamanis, Miltiadis and Brockschmidt, Marc},520  journal={arXiv preprint arXiv:1909.09436},521  year={2019}522}523 524### Contributions525 526Thanks to [@SBrandeis](https://github.com/SBrandeis) for adding this dataset.527