Team Ai
Modelpublic

chrisswillss98/dpo_mcqa_quantizedBitsAndBytes

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
0likes9downloads
Model Card

<!-- This model card has been generated automatically according to the information the Trainer had access to. You should probably proofread and complete it, then remove this comment. -->

mcqaquantBBQ

This model is a fine-tuned version of TinyLlama/TinyLlama-1.1B-Chat-v1.0 on an unknown dataset. It achieves the following results on the evaluation set:

  • —Loss: 0.9790
  • —Rewards/chosen: -0.9912
  • —Rewards/rejected: -0.9047
  • —Rewards/accuracies: 0.5
  • —Rewards/margins: -0.0865
  • —Logps/rejected: -26.9191
  • —Logps/chosen: -28.3753
  • —Logits/rejected: -3.3916
  • —Logits/chosen: -3.3931

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • —learning_rate: 1e-05
  • —trainbatchsize: 8
  • —evalbatchsize: 1
  • —seed: 42
  • —optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
  • —lrschedulertype: cosine
  • —lrschedulerwarmup_ratio: 0.1
  • —num_epochs: 10

Training results

Training LossEpochStepValidation LossRewards/chosenRewards/rejectedRewards/accuraciesRewards/marginsLogps/rejectedLogps/chosenLogits/rejectedLogits/chosen
0.66381.7241500.7089-0.02010.00540.3462-0.0254-17.8186-18.6642-3.5325-3.5335
0.56963.44831000.7434-0.2832-0.25590.4615-0.0273-20.4312-21.2954-3.4911-3.4921
0.37515.17241500.8633-0.5219-0.47770.4615-0.0443-22.6490-23.6828-3.4444-3.4462
0.21266.89662001.0177-0.7687-0.62090.3846-0.1478-24.0814-26.1501-3.4037-3.4053
0.17648.62072500.9790-0.9912-0.90470.5-0.0865-26.9191-28.3753-3.3916-3.3931

Framework versions

  • —PEFT 0.11.1
  • —Transformers 4.41.2
  • —Pytorch 2.3.1+cu118
  • —Datasets 2.20.0
  • —Tokenizers 0.19.1