Team Ai
Apppublic

jjderz/finetuning-language-models-tweets

sourceHugging Faceupdated 3y agoView on Hugging Face
0likes
App README

Google Sites Landing Page: https://sites.google.com/njit.edu/finetunedtoxiccommentmodel/home

Hugging Face: https://huggingface.co/spaces/jjderz/finetuning-language-models-tweets

Toxic Comment Classifier

This project is designed to identify various forms of toxicity in comments using a deep learning model built on the BERT architecture. It employs a pre-trained BERT model that is fine-tuned on a dataset of labeled Wikipedia comments. The project consists of two main components: training the model in toxic_trainer.ipynb and deploying the model as a web application using Streamlit in app.py.

toxic_trainer.ipynb

This Jupyter notebook contains the code for training a BERT-based model to perform multi-label classification of toxic comments. The primary elements of this notebook are:

  1. 1.Imports: Incorporating necessary libraries and modules.
  1. 1.Drive Mounting: Connecting to Google Drive to store the model and checkpoints.
  1. 1.Labels and dictionaries: Specifying labels and generating dictionaries for label-to-index and index-to-label mapping.
  1. 1.Dataset loading: Retrieving the dataset from a CSV file and tokenizing it using a pre-trained BERT tokenizer.
  1. 1.Preprocessing: Developing a function to preprocess the data by tokenizing text and encoding labels.
  1. 1.Model definition: Establishing the BERT-based model for multi-label classification.
  1. 1.Dataset encoding: Encoding the dataset using the preprocessing function.
  1. 1.Model training: Assembling and training the model using the encoded dataset.
  1. 1.Upload to Huggingface: Using Huggingface cli to push the model and tokenizer to the hub.
app.py

This Python script is utilized to launch the trained model as a web application with the help of the Streamlit library. The primary components of this script are:

  1. 1.Imports: Incorporating necessary libraries and modules.
  1. 1.Title: Exhibiting the title of the web application.
  1. 1.Input text area: Generating a text area for users to enter text for toxicity classification.
  1. 1.Model choice: Offering a selection for users to pick their preferred model for classification.
  1. 1.Tokenizer and model loading: Retrieving the tokenizer and model based on user selection.
  1. 1.Sentiment pipeline: Developing a sentiment analysis pipeline using the loaded tokenizer and model.
  1. 1.Submit button: Introducing a button for submitting the input text for classification.
  1. 1.Classification and result display: When the submit button is activated, the input text is classified using the sentiment pipeline, and the results are presented in a table.
Dependencies
  • —Python 3.6+
  • —TensorFlow 2.x
  • —transformers
  • —datasets
  • —streamlit
  • —pandas