Team Ai
Modelpublic

kohendru/distilbert-base-uncased-amazon-sentiment-analysis

sourceHugging Facemitupdated 2y agoView on Hugging Face
0likes20downloads
Model Card

distilbert-base-uncased-amazon-sentiment-analysis

Base Model

  • —BERT: BERT is a transformer-based model designed to pre-train deep bidirectional representations by jointly conditioning on both left and right context in all layers.
  • —DistilBERT: DistilBERT is a smaller, faster, and more efficient version of BERT. It uses knowledge distillation to reduce the model size by approximately 60% while retaining 97% of BERT’s language understanding capabilities.

Dataset

The dataset obtained from kaggle with title "Amazon Reviews for Sentiment Analysis" by Adam Bittlingmayer. The dataset contains columns "title," "text," and "label," with a total of 4,000,000 data entries (I only use 5% of the data for now).

Dataset Example

titletextlabel
Stuning even for the non-gamerThis sound track was beautiful! It paints the ...2
The best soundtrack ever to anything.I'm reading a lot of reviews saying that this ...2
Amazing!This soundtrack is my favorite music of all ti...2
Excellent SoundtrackI truly like this soundtrack and I enjoy video...2
Remember, Pull Your Jaw Off The Floor After He...If you've played the game, you know how divine...2
.........
Unbelievable- In a Bad WayWe bought this Thomas for our son who is a hug...1
Almost Great, Until it Broke...My son recieved this as a birthday gift 2 mont...1
Disappointed !!!I bought this toy for my son who loves the "Th...1
Classic Jessica MitfordThis is a compilation of a wide range of Mitfo...2
Comedy Scene, and Not HeardThis DVD will be a disappointment if you get i...1

Evaluation

When I try to train the model with a large number of epochs, it starts to overfit when the epoch reaches 6 or 7. So, I only use 5 epochs for this model.

EpochTraining LossValidation LossAccuracyPrecision MacroRecall MacroF1 Macro
10.1442000.1397920.9485750.9485710.9485830.948574
20.1244000.1456470.9516500.9518170.9517090.951649
30.1129000.1488250.9536000.9536030.9536160.953600
40.0812000.1551140.9539250.9539210.9539320.953924
50.1024000.1712980.9536000.9535980.9536120.953600
py
results = trainer.evaluate()
print(results)

"""
{
    'eval_accuracy': 0.953925, 
    'eval_precision_macro': 0.9539209871607255, 
    'eval_recall_macro': 0.9539319939428168, 
    'eval_f1_macro': 0.9539242719746999, 
    'eval_loss': 0.15511418879032135, 
    'eval_runtime': 90.9442, 
    'eval_samples_per_second': 439.83, 
    'eval_steps_per_second': 6.872, 
    'epoch': 5.0
}
"""

How to use the model?

py
from transformers import pipeline

model_name = "kohendru/distilbert-base-uncased-amazon-sentiment-analysis"
nlp = pipeline("text-classification", model=model_name, tokenizer=model_name)

reviews = [
    "I love this product! It works great and has exceeded my expectations.",
    "Worst purchase ever. Completely useless and waste of money.",
    "The product is okay, but could be improved in terms of quality.",
    "Amazing! Will definitely buy again."
]

for review in reviews:
    result = nlp(review)
    print(f"Review: {review}")
    print(f"Sentiment: {result[0]['label']}, Confidence: {result[0]['score']:.4f}")
    print("-" * 50)

"""
Review: I love this product! It works great and has exceeded my expectations.
Sentiment: Good Review, Confidence: 0.9950
--------------------------------------------------
Review: Worst purchase ever. Completely useless and waste of money.
Sentiment: Bad Review, Confidence: 0.9958
--------------------------------------------------
Review: The product is okay, but could be improved in terms of quality.
Sentiment: Bad Review, Confidence: 0.5947
--------------------------------------------------
Review: Amazing! Will definitely buy again.
Sentiment: Good Review, Confidence: 0.9942
--------------------------------------------------
"""