Team Ai
Apppublic

YaekobB/Toxic-Comment-Classification

sourceHugging Facemitupdated 5mo agoView on Hugging Face
0likes
App README

🧠 Toxic Comment Classification β€” Explainable Multi-Label NLP Model

<p align="center"> <img src="banner.png" alt="Toxic Comment Classification Banner" width="100%"> </p>

<p align="center"> <b>DistilBERT-based multi-label classifier for detecting toxic online comments with explainability powered by Captum Integrated Gradients (IG).</b> </p>


πŸš€ Overview

This project presents an explainable AI system for identifying toxic comments in text, built using a fine-tuned Transformer model (DistilBERT). It performs multi-label classification across six toxicity categories while offering token-level explanations for each prediction.

🧩 Labels

  • β€”toxic
  • β€”severe_toxic
  • β€”obscene
  • β€”threat
  • β€”insult
  • β€”identity_hate

🎯 Objectives

  • β€”Fine-tune DistilBERT for robust multi-label toxicity detection
  • β€”Enhance interpretability using Captum Integrated Gradients
  • β€”Deploy a real-time, user-friendly Gradio interface

πŸ§ͺ How to Use the Demo

  1. 1.Type or paste any comment in the text box
  2. 2.Click β€œClassify” to view per-label probabilities and predictions
  3. 3.Open the β€œExplain” tab β†’ select a target label
  4. 4.Generate a heatmap showing which words support (red) or oppose (blue) the decision

🧠 Example Inputs

ExampleExpected Labels
β€œYou are a complete idiot.”toxic / insult
β€œI will kill you tomorrow.”threat / toxic
β€œThanks for your help today!”non-toxic
β€œGo away, you people don’t belong here.”identity_hate / insult

βš™οΈ Technical Stack

ComponentTechnology
Language ModelDistilBERT (distilbert-base-uncased)
FrameworksPyTorch β€’ Transformers β€’ Gradio
ExplainabilityCaptum (Integrated Gradients)
TrainingStratified splits β€’ Early Stopping β€’ Regularization
VisualizationGradio UI + Captum HTML heatmaps
DeploymentHugging Face Spaces

πŸ“‚ Project Structure

.
β”œβ”€β”€ app.py                # Gradio app entry point
β”œβ”€β”€ requirements.txt      # Runtime dependencies
β”œβ”€β”€ artifacts/
β”‚   β”œβ”€β”€ best/             # Fine-tuned model weights + tokenizer
β”‚   └── thresholds.json   # Tuned thresholds for each label
└── README.md             # (this file)

πŸ“Š Model Training Summary

  • β€”Dataset: Jigsaw Toxic Comment Classification Challenge
  • β€”Tokenization: DistilBERT (max length = 256)
  • β€”Loss: Binary Cross-Entropy with Logits (BCEWithLogitsLoss)
  • β€”Optimizer: AdamW (learning rate = 2e-5, weight decay = 0.02)
  • β€”Regularization: Dropout (head=0.5, encoder=0.2)
  • β€”Evaluation Metrics: Macro F1 β€’ Precision β€’ Recall β€’ AUC
  • β€”Explainability: Captum Layer Integrated Gradients (LIG)

πŸ–₯️ Live Demo

πŸš€ Try the interactive demo on Hugging Face Spaces: πŸ”— [yaekobB / Toxic-Comment-Classification](https://huggingface.co/spaces/yaekobB/Toxic-Comment-Classification)

🧰 Dependencies

txt
transformers>=4.41.0
torch>=2.2.0
safetensors>=0.4.2
gradio>=4.20.0
captum>=0.7.0
pandas>=2.0.0
numpy>=1.24.0


πŸͺͺ License

This project is licensed under the MIT License. You are free to use, modify, and distribute this work with attribution.


<p align="center"> <i>β€œBuilding safer and explainable AI for online interactions.”</i> </p>