Team Ai
Modelpublic

rajaykumar12959/python-codesearchnet-tokenizer

sourceHugging Facemitupdated 10mo agoView on Hugging Face
1likes
README.md64 linesDownload Raw Back to root
1---2language: en3license: mit4tags:5- tokenizer6- python7- code-search-net8- bpe9library_name: transformers10base_model: gpt211---12 13# Tokenizer for Python Code (Trained on CodeSearchNet)14 15## Model Description16 17This is a custom Byte-Pair Encoding (BPE) tokenizer, initialized from a `gpt2` tokenizer and further trained on the Python subset of the [CodeSearchNet dataset](https://huggingface.co/datasets/claudios/code_search_net). The tokenizer is designed to efficiently tokenize Python code, which can be useful for various downstream tasks like code generation, code completion, and code analysis.18 19## Training Data20 21The tokenizer was trained on the `whole_func_string` column of the `train` split from the `claudios/code_search_net` dataset, specifically focusing on Python code examples. The training corpus consisted of approximately 412,178 Python function strings.22 23## Training Procedure24 251.  **Base Tokenizer**: Started with a pre-trained `gpt2` tokenizer.262.  **Training**: The `train_new_from_iterator` method from `transformers.PreTrainedTokenizerFast` was used to train a new vocabulary and merges from the `CodeSearchNet` Python code corpus. The new vocabulary size was set to 52,000 tokens.27 28## How to Use29 30You can load and use this tokenizer with the `transformers` library:31 32```python33from transformers import AutoTokenizer34 35# Load the tokenizer from the Hugging Face Hub36tokenizer = AutoTokenizer.from_pretrained("rajaykumar12959/new_tokeniser")37 38# Example usage39example_code = """class LinearLayer():40    def __init__(self, input_size, output_size):41        self.weight = torch.randn(input_size, output_size)42        self.bias = torch.zeros(output_size)43 44    def __call__(self, x):45        return x @ self.weights + self.bias46    """47 48tokens = tokenizer.tokenize(example_code)49print(tokens)50# Output will be similar to:51# ['class', 'ĠLinear', 'Layer', '():', 'ĊĠĠĠ', 'Ġdef', 'Ġ__', 'init', '__(', 'self', ',', 'Ġinput', '_', 'size', ',', 'Ġoutput', '_', 'size', '):', 'ĊĠĠĠĠĠĠĠ', 'Ġself', '.', 'weight', 'Ġ=', 'Ġtorch', '.', 'randn', '(', 'input', '_', 'size', ',', 'Ġoutput', '_', 'size', ')', 'ĊĠĠĠĠĠĠĠ', 'Ġself', '.', 'bias', 'Ġ=', 'Ġtorch', '.', 'zeros', '(', 'output', '_', 'size', ')', 'ĊĊĠĠĠ', 'Ġdef', 'Ġ__', 'call', '__(', 'self', ',', 'Ġx', '):', 'ĊĠĠĠĠĠĠĠ', 'Ġreturn', 'Ġx', 'Ġ@', 'Ġself', '.', 'weights', 'Ġ+', 'Ġself', '.', 'bias', 'ĊĠĠĠĠ']52 53encoded_input = tokenizer(example_code, return_tensors="pt")54print(encoded_input)55```56 57## License58 59This tokenizer is licensed under the MIT License.60 61## Author62 63[rajaykumar12959](https://huggingface.co/rajaykumar12959)64