Team Ai
Modelpublic

Aukrk/MLOPS_group-v4

sourceHugging Faceapache-2.0updated 4mo agoView on Hugging Face
0likes14downloads
Model Card

license: apache-2.0 libraryname: transformers pipelinetag: text-classification tags:

  • —transformers
  • —text-classification
  • —distilbert
  • —sms-spam
  • —mlops
  • —kaggle
  • —wandb
  • —huggingface
  • —iit-jodhpur
  • —group-36
  • —G25AIT2016 datasets:
  • —ucirvine/smsspam basemodel: distilbert-base-uncased language:
  • —en ---

MLOPS_group-v4 — SMS Spam Classification

This repository is part of the MLOps Group 36 Project for the PGD AI Programme, IIT Jodhpur.

The project implements an end-to-end MLOps pipeline for SMS spam classification using DistilBERT, with GitHub, Kaggle, Weights & Biases, Hugging Face Hub, Docker, and GitHub Actions.

Contributor

Anu Kumar Roll Number: G25AIT2016

Project Links

ResourceLink
GitHub Repositoryhttps://github.com/g25ait2032-prog/mlops-group36-iitj
Kaggle Notebook - G25AIT2016https://www.kaggle.com/code/anukumarkg25ait2016/mlops-group36-data-preprocessing-g25ait2016
W&B Run - G25AIT2016https://wandb.ai/g25ait2032-iit-jodhpur/MLOPS_Group/runs/j5fk4zll
W&B Project Dashboardhttps://wandb.ai/g25ait2032-iit-jodhpur/MLOPS_Group
Hugging Face Modelhttps://huggingface.co/Aukrk/MLOPS_group-v4

Model Details

ItemValue
Base modeldistilbert-base-uncased
TaskBinary text classification
Classesham, spam
DatasetUCI SMS Spam Collection
FrameworkHugging Face Transformers
Output labels0 = ham, 1 = spam

Contribution Summary

This repository is linked to the G25AIT2016 Task 2 workflow.

The completed contribution includes:

  • —Loading the UCI SMS Spam dataset
  • —Cleaning and normalising SMS text
  • —Removing missing and duplicate records
  • —Creating stratified train, validation, and test splits
  • —Creating id2label.json and label2id.json
  • —Running data sanity checks
  • —Logging data-preparation metrics to W&B
  • —Publishing this Hugging Face model repository for project traceability

Dataset Preparation Summary

MetricValue
Raw samples5,574
Duplicates removed415
Cleaned samples5,159
Train rows3,611
Validation rows774
Test rows774
Sanity checks passed21 / 21
Leakage checkPassed

Label Mapping

json
{
  "0": "ham",
  "1": "spam"
}

How to Use

python
from transformers import pipeline

classifier = pipeline(
    "text-classification",
    model="Aukrk/MLOPS_group-v4"
)

text = "Congratulations! You have won a free iPhone. Click here now."
result = classifier(text)

print(result)

Load Model Directly

python
from transformers import AutoTokenizer, AutoModelForSequenceClassification

tokenizer = AutoTokenizer.from_pretrained("Aukrk/MLOPS_group-v4")
model = AutoModelForSequenceClassification.from_pretrained("Aukrk/MLOPS_group-v4")

Example Inputs

TextExpected Output
Congratulations! You have won a free prize. Click here now.spam
Can we meet tomorrow at 5 PM?ham

W&B Traceability

The G25AIT2016 W&B run records data-preparation metrics such as:

  • —Raw sample count
  • —Duplicate removal count
  • —Cleaned sample count
  • —Train / validation / test split sizes
  • —Sanity check status
  • —Leakage check status

W&B Run: https://wandb.ai/g25ait2032-iit-jodhpur/MLOPS_Group/runs/j5fk4zll

Model Context

This model repository is published under the G25AIT2016 Hugging Face account for Group 36 project traceability.

The model artefact follows the Group 36 DistilBERT SMS spam classification workflow and is linked with the data-preparation contribution completed by Anu Kumar - G25AIT2016.

Limitations

  • —The dataset is relatively small and focused on SMS messages.
  • —The model may not generalise well to long emails, non-English messages, or modern scam formats.
  • —Boundary cases mixing normal conversation and promotional text may be misclassified.
  • —This is an academic MLOps demonstration and should not be used as the only spam detection control in production.

Intended Use

This repository is intended for:

  • —Academic MLOps demonstration
  • —SMS spam classification testing
  • —Hugging Face deployment evidence
  • —W&B traceability evidence
  • —GitHub Actions / Docker inference integration

Not Intended For

  • —Production-grade fraud detection
  • —Legal, financial, or safety-critical filtering
  • —Detecting all phishing or scam variants without further validation

Authors

MLOps Group 36 PGD AI Programme, IIT Jodhpur

Contributor for this repository: Anu Kumar - G25AIT2016