Team Ai
Modelpublic

SaikatM/Code-Gemma-v1

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
0likes22downloads
Model Card

Model Card for Model ID

<!-- Provide a quick summary of what the model is/does. --> This model is a fine-tuned version of google/gemma-2b on an SaikatM/Code-Platypus dataset.

Model Details

Model Description

  • —Finetuned from model: google/gemma-2b

Model Sources

Training code can be found here: https://github.com/Saikat-M/LLM-Finetuning

Direct Use

  • —Code generation tasks

Training Data

  • —Dataset: https://huggingface.co/datasets/SaikatM/Code-Platypus
  • —Source Dataset: https://huggingface.co/datasets/garage-bAInd/Open-Platypus

Training Procedure

Used QLoRA from PEFT and used SFTTrainer.

Preprocessing

From the Open-Platypus dataset filtering-out rows which has leetcodene in it's datasource column.

Training Hyperparameters

LoraConfig(

r=4, loraalpha=2, targetmodules=modules, loradropout=0.05, bias="none", tasktype="CAUSAL_LM" )

TrainingArguments(

outputdir="gemma-2b-code-platypus", numtrainepochs=1, perdevicetrainbatchsize=4, gradientaccumulationsteps=4, gradientcheckpointing=True, optim="pagedadamw8bit", loggingsteps=1, savestrategy="epoch", bf16=False, tf32=False, learningrate=2e-4, maxsteps= 100, maxgradnorm=0.3, warmupratio=0.03, lrschedulertype="constant", pushtohub=False, reportto="tensorboard", )

SFTTrainer(

model=model, traindataset=traindata, evaldataset=testdata, datasettextfield="text", peftconfig=loraconfig, maxseqlength=512, tokenizer=tokenizer, args=training_arguments, )

Speeds, Sizes, Times

Took around 1 hour to train.

Results

  • —Test Result 1:
Write a fucntion to sort a list in python

Answer:

def sort_list(list):
    return sorted(list)<eos>
Response:  None
  • —Test Result 2:
Write a function to count Consonants in a Given Word in Python

Response:  None
  • —Test Result 3:
Write a function to count the number of vowels in a given string in Python.

Example 1:

Input:  s =  "leetcodeisgreat"
Output:  5
Explanation:  The vowels are  'e', 'i', 'a', 'o', and 'u'.

Example 2:

Input:  s =  "leetcodeisgreat"
Output:  0
Explanation:  The vowels are  'e', 'i', 'a', 'o', and 'u'.

Constraints:

*   1 <= s.length <= 100
*   s consists of lowercase English letters.


def countVowels(s):
    count = 0
    for c in s:
        if c in 'aeiou':
            count += 1
    return count
<eos>
Response:  None

Compute Infrastructure

Trained in Google Colab

Hardware

T4 GPU Hardware accelerator.