SaikatM/Code-Gemma-v1
Model Card for Model ID
<!-- Provide a quick summary of what the model is/does. --> This model is a fine-tuned version of google/gemma-2b on an SaikatM/Code-Platypus dataset.
Model Details
Model Description
- Finetuned from model: google/gemma-2b
Model Sources
Training code can be found here: https://github.com/Saikat-M/LLM-Finetuning
Direct Use
- Code generation tasks
Training Data
- Dataset: https://huggingface.co/datasets/SaikatM/Code-Platypus
- Source Dataset: https://huggingface.co/datasets/garage-bAInd/Open-Platypus
Training Procedure
Used QLoRA from PEFT and used SFTTrainer.
Preprocessing
From the Open-Platypus dataset filtering-out rows which has leetcodene in it's datasource column.
Training Hyperparameters
LoraConfig(
r=4, loraalpha=2, targetmodules=modules, loradropout=0.05, bias="none", tasktype="CAUSAL_LM" )
TrainingArguments(
outputdir="gemma-2b-code-platypus", numtrainepochs=1, perdevicetrainbatchsize=4, gradientaccumulationsteps=4, gradientcheckpointing=True, optim="pagedadamw8bit", loggingsteps=1, savestrategy="epoch", bf16=False, tf32=False, learningrate=2e-4, maxsteps= 100, maxgradnorm=0.3, warmupratio=0.03, lrschedulertype="constant", pushtohub=False, reportto="tensorboard", )
SFTTrainer(
model=model, traindataset=traindata, evaldataset=testdata, datasettextfield="text", peftconfig=loraconfig, maxseqlength=512, tokenizer=tokenizer, args=training_arguments, )
Speeds, Sizes, Times
Took around 1 hour to train.
Results
- Test Result 1:
Write a fucntion to sort a list in python
Answer:
def sort_list(list):
return sorted(list)<eos>
Response: None- Test Result 2:
Write a function to count Consonants in a Given Word in Python
Response: None- Test Result 3:
Write a function to count the number of vowels in a given string in Python.
Example 1:
Input: s = "leetcodeisgreat"
Output: 5
Explanation: The vowels are 'e', 'i', 'a', 'o', and 'u'.
Example 2:
Input: s = "leetcodeisgreat"
Output: 0
Explanation: The vowels are 'e', 'i', 'a', 'o', and 'u'.
Constraints:
* 1 <= s.length <= 100
* s consists of lowercase English letters.
def countVowels(s):
count = 0
for c in s:
if c in 'aeiou':
count += 1
return count
<eos>
Response: NoneCompute Infrastructure
Trained in Google Colab
Hardware
T4 GPU Hardware accelerator.
