Team Ai
Modelpublic

Woolv7007/egyptian-text-classification

sourceHugging Facemitupdated 1y agoView on Hugging Face
1likes29downloads
Model Card

Egyptian Dialect Text Classification by Fine-tuning AraBERT ๐Ÿ‡ช๐Ÿ‡ฌ

This model is a fine-tuned version of aubmindlab/bert-base-arabertv2, specifically trained on Egyptian Arabic text for hate speech and offensive language classification

with 92% accurency.

It classifies text into 6 categories :

CategoryLabel
OffensiveOffensive text
NeutralNeutral text
RacismRacism
SexismSexism
Religious DiscriminationReligious discrimination
AdsAdvertisements

Training Details

  • โ€”Base Model: AraBERT v2 (aubmindlab/bert-base-arabertv2)
  • โ€”Dataset: Egyptian Arabic hate speech dataset with 6 labeled categories
  • โ€”Epochs: Trained for 30 epochs with early stopping after 5 epochs without improvement
  • โ€”Optimizer Enhancements: Used label smoothing with a factor of 0.1 for better generalization
  • โ€”Best Model Selection: Based on weighted average F1-score

Performance Metrics

Training was done using a GPU (cuda). Below is a snapshot of model performance across the first 15 epochs:

EpochTrain LossVal LossAccuracyPrecisionRecallF1
11.68911.52270.49800.50270.49800.4741
21.19530.94910.77840.78830.77840.7758
30.76120.70100.86700.86930.86700.8673
40.62650.63630.90350.90420.90350.9031
50.55050.65470.89960.89950.89960.8990
60.51190.68610.89310.90180.89310.8947
70.47790.66750.90480.90660.90480.9052
80.46730.63530.92180.92380.92180.9222
90.45420.66140.91260.91360.91260.9125
100.44440.66180.92310.92380.92310.9233
110.43590.66890.92310.92350.92310.9230
120.43440.71200.90610.90970.90610.9067
130.43250.72480.90610.91050.90610.9068
140.43690.69460.91790.92210.91790.9189
150.42890.68640.91530.91710.91530.9157

Using the Model Without a Token (For End Users)

Since the model Woolv7007/egyptian-text-classification and the labels.json file are publicly available, you can load the model, tokenizer, and labels directly without needing any Hugging Face token or special setup.

How to automatically load and use the model with labels:

  • โ€”Download the labels.json file directly from Hugging Face Hub.
  • โ€”Load the model and tokenizer without a token.
  • โ€”Perform prediction on your texts and map the predicted class index to its label from the labels list.
  • โ€”Or you can call the labels.json file address with this code
labels_url = f"https://huggingface.co/{model_name}/resolve/main/labels.json"
labels = requests.get(labels_url).json()

Example Python Code:

python
import requests
import torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification

# Model name on Hugging Face Hub
model_name = "Woolv7007/egyptian-text-classification"

# Load labels.json from the public repository without a token
labels_url = f"https://huggingface.co/{model_name}/resolve/main/labels.json"
labels = requests.get(labels_url).json()

# Load model and tokenizer without a token
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForSequenceClassification.from_pretrained(model_name)
model.eval()

# Simple prediction function that returns the predicted label
def predict(text):
    inputs = tokenizer(text, return_tensors="pt", truncation=True, padding=True, max_length=256)
    with torch.no_grad():
        outputs = model(**inputs)
        pred_id = torch.argmax(outputs.logits, dim=1).item()
    return labels[pred_id]

# Verbose prediction with probability scores for each label (optional)
def predict_verbose(text):
    inputs = tokenizer(text, return_tensors="pt", truncation=True, padding=True, max_length=256)
    with torch.no_grad():
        logits = model(**inputs).logits
        probs = torch.softmax(logits, dim=1).squeeze().tolist()
    for label, prob in zip(labels, probs):
        print(f"{label}: {prob:.2%}")
    return labels[torch.argmax(logits).item()]

# Example usage
text = "ุทุฒ ููŠ ุงูŠ ุญุฏ ู…ุด ุนุงุฌุจู‡ ุดุบู„ูŠ"
print("Text:", text)
print("Predicted label:", predict(text))

# To see detailed output with probabilities, uncomment:
#print("Prediction with probabilities:")
#predict_verbose(text)