google/code_x_glue_ct_code_to_text
Dataset Card for "code_x_glue_ct_code_to_text" Dataset Summary CodeXGLUE code-to-text dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Text/code-to-text The dataset we use comes from CodeSearchNet and we filter the dataset as the following: Remove examples that codes cannot be parsed into an abstract syntax tree. Remove examples that #tokens of documents is < 3 or >256 Remove examples that documents contain special tokens (e.g. <img… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_ct_code_to_text.
Convert dataset to Parquet (#5)
Delete legacy JSON metadata (#4)
rename configs to config_name
Update CodeSearchNet data file URLs (#3)
Reorder split names (#1)
add dataset_info in dataset metadata
remove dummmy data
fix task_ids
Align/fix license metadata info (#4613)
Remove config names as yaml keys (#4367)
Update datasets task tags to align tags with models (#4067)
Update files from the datasets library (from 1.18.0)
Update files from the datasets library (from 1.8.0)
