NovelAI/genji-python-6B
42318
1---2language:3- en4tags:5- pytorch6- causal-lm7license: apache-2.08datasets:9- the Pile10---11 12# Genji-python 6B13 14For example usage or to easily use the model you can check our colab notebook:15[Notebook](https://colab.research.google.com/drive/1PnWpx02IEUkY8jhLKd_NewUGEXahAska?usp=sharing)16 17## Model Description18 19Genji is a transformer model finetuned on EleutherAI's GPT-J 6B model. This particular model is trained on python only code approaching 4GB in size.20 21| Hyperparameter | Value | 22|-------------------|--------|23| n_parameters | 6,053,381,344 |24| n_layers | 28* |25| d_model | 4,096 |26| d_ff | 16,384 |27| n_heads | 16 |28| d_head | 256 |29| n_ctx | 2,048 |30| n_vocab | 50,400 (same tokenizer as GPT-2/3) |31| position encoding | [Rotary position encodings (RoPE)](https://arxiv.org/abs/2104.09864) |32| RoPE dimensions | [64](https://github.com/kingoflolz/mesh-transformer-jax/blob/f2aa66e0925de6593dcbb70e72399b97b4130482/mesh_transformer/layers.py#L223) |33 34`*` each layer consists of one feedforward block and one self attention block35 36The model consists of 28 layers with a model dimension of 4096, and a feedforward dimension of 16384. The model37dimension is split into 16 heads, each with a dimension of 256. Rotary position encodings (RoPE) was applied to 6438dimensions of each head. The model is trained with a tokenization vocabulary of 50257, using the same set of BPEs as39GPT-2/GPT-3.40 41## Training data42 43GPT-J 6B was pretrained on the [Pile](pile.eleuther.ai), a large scale curated dataset created by EleutherAI for the purpose of training this model. After the pre-training, it's finetuned on the python code that was taken from the Pile.44 45## Training procedure46 47Genji-python-6B is trained for 20k steps on around 655 million tokens with learning rate of 2e-0648 49## Intended Use50 51This model is trained for assistence on writing python code and having fun trying weird stuff with it. 52 53### How to use54 55This model is only usable with our fork because GPT-J is not merged to the main transformers repo yet. When it's merged, we will make this model easily loadable.56For now, you need to use this fork:57[Fork](https://github.com/finetuneanon/transformers)58 59to install with pip:60```bash61pip install git+https://github.com/finetuneanon/transformers@gpt-neo-localattention3-rp-b62```63 64This model takes more than 16 gigs of RAM to load. If you want more efficient and faster loading, please check our split model.65We recommend the usage of the model as FP16. That way, it fits in 16GB VRAM cards.66 67How to use:68```python69from transformers import (70 AutoTokenizer,71 AutoModelForCausalLM,72 GPTNeoForCausalLM,73)74 75model = AutoModelForCausalLM.from_pretrained("NovelAI/genji-python-6B", use_auth_token=True).half().eval().cuda()76tokenizer = AutoTokenizer.from_pretrained("EleutherAI/gpt-neo-2.7B")77 78text = '''def print_customer_name'''79 80tokens = tokenizer(text, return_tensors="pt").input_ids81generated_tokens = model.generate(tokens.long().cuda(), use_cache=True, do_sample=True, top_k=50, temperature=0.3, top_p=0.9, repetition_penalty=1.125, min_length=1, max_length=len(tokens[0]) + 400, pad_token_id=tokenizer.eos_token_id)82last_tokens = generated_tokens[0][len(tokens[0]):]83generated_text = tokenizer.decode(last_tokens)84print("Generation:\n" + generated_text)85```86When ran, this code generates:87```python88Prompt:89def print_customer_name90Generation:91(self, customer):92 """Print the name of a customer."""93 if not self.is_valid():94 return95 96 print("Customer: {}".format(customer))97```98 99For example usage, you can see our colab notebook as well:100[Notebook](https://colab.research.google.com/drive/1PnWpx02IEUkY8jhLKd_NewUGEXahAska?usp=sharing)101 102## Eval results103 104TBD105 106## Acknowledgements107 108This project was possible because of the compute provided by the109[TPU Research Cloud](https://sites.research.google/trc/)110 111and [EleutherAI](https://eleuther.ai/) for pretraining of the GPT-J 6B.112 113Thanks to everyone who contributed to this project!114 115- [Aero](https://github.com/AeroScripts)116- [Finetune](https://github.com/finetuneanon)117- [Kurumuz](https://github.com/kurumuz)