datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CodeSearchNet-Python-LDUcodesearchnet-challenge-extended
CodeSearchNet Challenge, extended: every search over every function
The CodeSearchNet Challenge (Husain et al., 2019) has
99 natural-language code searches, and experts rated a few candidate functions for each. This
dataset treats every rated function of a language as one codebase and searches all of it: for each
search, every function in its language is a candidate. The experts' ratings are kept, and the
pairs they never rated but a search tool returned were rated on the same… See the full description on the dataset page: https://huggingface.co/datasets/Scoolar/codesearchnet-challenge-extended.codesearchnet-python-linelen40-fullcodesearchnet-python-rebalancedcodesearchnet-python-pep8-fullcodesearchnet-python-rebalanced-linelevel-scoredcodesearchnet-py-linelen40-rebalanced200k-v1code_search_net_python_filtered_top50k
Dataset Card for "code_search_net_python_filtered_top50k"
More Information needed
code_search_net_python_processed_400k
Dataset Card for "code_search_net_python_processed_400k"
More Information needed
codesearchnet-defect-labelscode_search_net_filtered_top100
Dataset Card for "code_search_net_filtered_top100"
More Information needed
code_searchnet_reduced
Dataset Card for "code_searchnet_reduced"
More Information needed
code_searchnet_reduced_train
Dataset Card for "code_searchnet_reduced_train"
More Information needed
code_searchnet_reduced_val
Dataset Card for "code_searchnet_reduced_val"
More Information needed
code_search_net_filtered_34kFiltered version of code search net python subset, with filtering based on perplexity with/without docstring, learning value/quality classifiers, and manual filtering.
Original data with perplexity filtering is from here, with credit to bjoernp.
