Team Ai
Modelpublic

Iltaf/sst2-lora-distilbert

sourceHugging Faceupdated 22d agoView on Hugging Face
0likes67downloads
Model Card

SST-2 LoRA adapter on distilbert-base-uncased-finetuned-sst-2-english

A LoRA adapter fine-tuned on a subset of SST-2 (binary sentiment) using free Google Colab hardware. This is the deliverable for the E3 project: fine-tune a small open model with LoRA, measure a baseline before training, train, measure again, and report the change honestly even if small.

What this adapter is

  • —Base model: distilbert-base-uncased-finetuned-sst-2-english (already fine-tuned on full SST-2 by HuggingFace).
  • —Method: LoRA via peft - r=8, alpha=16, dropout=0.05, targets q_lin, v_lin, and (via modules_to_save) the classifier and pre_classifier heads.
  • —Training data: 8,000 rows sampled from nyu-mll/glue SST-2 train, seeded with SEED=42.
  • —Epochs: 3 - Batch: 32 - LR: 2e-4 - Precision: fp16
  • —Hardware: free Google Colab T4 (15 GB)
  • —Runtime: ~36 seconds of training (750 steps).

Results (honest before / after)

Same evaluation code, same tokenizer, same collator, same test/val splits for both the baseline and the LoRA-trained model. The only thing that changed is the model weights.

SplitnBaseline accAfter accDelta accBaseline F1After F1Delta F1
test720.90280.9028+0.00000.90190.9019+0.0000
val8000.91130.9087-0.00250.91120.9087-0.0025
full SST-2 val8720.90830.9083+0.00000.90810.9081+0.0000

Headline

The LoRA adapter did not improve over the frozen base checkpoint. On the 72-row test set and on the full 872-row SST-2 validation set, predictions are identical. On the 800-row in-training validation set, the adapter is 0.50 points worse (accuracy and macro-F1).

Training loss fell from 0.050 (epoch 1) to 0.037 (epoch 3) while validation loss did not improve - the classic signature of mild overfitting on a small training subset with an already-strong base model. There was very little room for a small adapter to improve on a checkpoint that was already fine-tuned on the full SST-2.

This is a negative result and is reported as such. It is the intended finding of the E3 exercise: the lesson is the method (baseline -> train -> measure -> report), not the magnitude of the delta.

Training curve

EpochTrain lossVal lossVal accuracyVal F1
10.05050.33360.90630.9061
20.03800.35890.90500.9049
30.03680.33610.90380.9036

The checkpoint reported above is epoch 1, selected by load_best_model_at_end=True on validation macro-F1.

Limitations (read these before using the adapter)

  1. 1.The test set has 72 rows. One flipped prediction is +/-1.39% accuracy. Treat any single test number as noisy.
  2. 2.The training set is 8,000 rows, about 12% of SST-2 train. This was a deliberate scope limit to fit free Colab.
  3. 3.The base model was already fine-tuned on full SST-2, so the headroom for improvement was small from the start.
  4. 4.No hyperparameter search. r=8, alpha=16, lr=2e-4, 3 epochs were chosen as standard defaults, not tuned.
  5. 5.modulestosave includes the classifier head, so part of the (non-)improvement is due to head training, not purely LoRA on the attention projections.

Reproduce

python
from transformers import AutoTokenizer, AutoModelForSequenceClassification
from peft import PeftModel

BASE = "distilbert-base-uncased-finetuned-sst-2-english"
ADAPTER = "Iltaf/sst2-lora-distilbert"

tok = AutoTokenizer.from_pretrained(ADAPTER)
base = AutoModelForSequenceClassification.from_pretrained(BASE)
model = PeftModel.from_pretrained(base, ADAPTER)

inputs = tok("this movie was a delight", return_tensors="pt")
logits = model(**inputs).logits
print("positive" if logits.argmax(-1).item() == 1 else "negative")

Provenance

  • —Notebook, seeds, and full trace available on request.
  • —SEED = 42 throughout (data split, training, and evaluation).
  • —Dataset: nyu-mll/glue, config sst2 (bare glue alias is deprecated in datasets>=4).

Reproducibility verification

Re-evaluated in a fresh Google Colab T4 runtime, loading base model and adapter directly from the HuggingFace Hub (no training code, no cached state):

SplitnReportedReproducedMatch
test720.9028 / 0.9019 (acc / macro-F1)0.9028 / 0.9019yes

The 72-row test set is rebuilt deterministically from nyu-mll/glue sst2 validation with SEED=42, consuming the RNG stream in the same order as training (train_idx draw first, then val_idx). This draw order matters: an earlier version of the check drew val_idx first and produced a different 72-row split, which scored 0.8750 for both the base model and the adapter. The corrected check reproduces the reported numbers exactly.