Team Ai
Datasetpublic

ytzi/the-stack-dedup-python-filtered-docstrings

This is a dataset originated from bigcode/the-stack-dedup with some filters applied. The filters filtered in this dataset are: remove_function_no_docstring remove_class_no_docstring remove_delete_markers

sourceHugging Faceupdated 3y agoView on Hugging Face
0likes92downloads
README.md89 linesDownload Raw Back to root
1---2dataset_info:3  config_name: main4  features:5  - name: hexsha6    dtype: string7  - name: size8    dtype: int649  - name: ext10    dtype: string11  - name: lang12    dtype: string13  - name: max_stars_repo_path14    dtype: string15  - name: max_stars_repo_name16    dtype: string17  - name: max_stars_repo_head_hexsha18    dtype: string19  - name: max_stars_repo_licenses20    sequence: string21  - name: max_stars_count22    dtype: int6423  - name: max_stars_repo_stars_event_min_datetime24    dtype: string25  - name: max_stars_repo_stars_event_max_datetime26    dtype: string27  - name: max_issues_repo_path28    dtype: string29  - name: max_issues_repo_name30    dtype: string31  - name: max_issues_repo_head_hexsha32    dtype: string33  - name: max_issues_repo_licenses34    sequence: string35  - name: max_issues_count36    dtype: int6437  - name: max_issues_repo_issues_event_min_datetime38    dtype: string39  - name: max_issues_repo_issues_event_max_datetime40    dtype: string41  - name: max_forks_repo_path42    dtype: string43  - name: max_forks_repo_name44    dtype: string45  - name: max_forks_repo_head_hexsha46    dtype: string47  - name: max_forks_repo_licenses48    sequence: string49  - name: max_forks_count50    dtype: int6451  - name: max_forks_repo_forks_event_min_datetime52    dtype: string53  - name: max_forks_repo_forks_event_max_datetime54    dtype: string55  - name: content56    dtype: string57  - name: avg_line_length58    dtype: float6459  - name: max_line_length60    dtype: int6461  - name: alphanum_fraction62    dtype: float6463  - name: original_content64    dtype: string65  - name: filtered:remove_function_no_docstring66    dtype: int6467  - name: filtered:remove_class_no_docstring68    dtype: int6469  - name: filtered:remove_delete_markers70    dtype: int6471  splits:72  - name: train73    num_bytes: 10299387742074    num_examples: 1276018275  download_size: 4101137519476  dataset_size: 10299387742077configs:78- config_name: main79  data_files:80  - split: train81    path: main/train-*82---83 84This is a dataset originated from [bigcode/the-stack-dedup](https://huggingface.co/datasets/bigcode/the-stack-dedup) with some filters applied.85The filters filtered in this dataset are:86 87 -  remove_function_no_docstring88 -  remove_class_no_docstring89 -  remove_delete_markers