YaekobB/Toxic-Comment-Classification
π§ Toxic Comment Classification β Explainable Multi-Label NLP Model
<p align="center"> <img src="banner.png" alt="Toxic Comment Classification Banner" width="100%"> </p>
<p align="center"> <b>DistilBERT-based multi-label classifier for detecting toxic online comments with explainability powered by Captum Integrated Gradients (IG).</b> </p>
π Overview
This project presents an explainable AI system for identifying toxic comments in text, built using a fine-tuned Transformer model (DistilBERT). It performs multi-label classification across six toxicity categories while offering token-level explanations for each prediction.
π§© Labels
- toxic
- severe_toxic
- obscene
- threat
- insult
- identity_hate
π― Objectives
- Fine-tune DistilBERT for robust multi-label toxicity detection
- Enhance interpretability using Captum Integrated Gradients
- Deploy a real-time, user-friendly Gradio interface
π§ͺ How to Use the Demo
- Type or paste any comment in the text box
- Click βClassifyβ to view per-label probabilities and predictions
- Open the βExplainβ tab β select a target label
- Generate a heatmap showing which words support (red) or oppose (blue) the decision
π§ Example Inputs
βοΈ Technical Stack
π Project Structure
.
βββ app.py # Gradio app entry point
βββ requirements.txt # Runtime dependencies
βββ artifacts/
β βββ best/ # Fine-tuned model weights + tokenizer
β βββ thresholds.json # Tuned thresholds for each label
βββ README.md # (this file)π Model Training Summary
- Dataset: Jigsaw Toxic Comment Classification Challenge
- Tokenization: DistilBERT (max length = 256)
- Loss: Binary Cross-Entropy with Logits (BCEWithLogitsLoss)
- Optimizer: AdamW (learning rate = 2e-5, weight decay = 0.02)
- Regularization: Dropout (head=0.5, encoder=0.2)
- Evaluation Metrics: Macro F1 β’ Precision β’ Recall β’ AUC
- Explainability: Captum Layer Integrated Gradients (LIG)
π₯οΈ Live Demo
π Try the interactive demo on Hugging Face Spaces: π [yaekobB / Toxic-Comment-Classification](https://huggingface.co/spaces/yaekobB/Toxic-Comment-Classification)
π§° Dependencies
transformers>=4.41.0
torch>=2.2.0
safetensors>=0.4.2
gradio>=4.20.0
captum>=0.7.0
pandas>=2.0.0
numpy>=1.24.0πͺͺ License
This project is licensed under the MIT License. You are free to use, modify, and distribute this work with attribution.
<p align="center"> <i>βBuilding safer and explainable AI for online interactions.β</i> </p>
