jjderz/finetuning-language-models-tweets
Google Sites Landing Page: https://sites.google.com/njit.edu/finetunedtoxiccommentmodel/home
Hugging Face: https://huggingface.co/spaces/jjderz/finetuning-language-models-tweets
Toxic Comment Classifier
This project is designed to identify various forms of toxicity in comments using a deep learning model built on the BERT architecture. It employs a pre-trained BERT model that is fine-tuned on a dataset of labeled Wikipedia comments. The project consists of two main components: training the model in toxic_trainer.ipynb and deploying the model as a web application using Streamlit in app.py.
toxic_trainer.ipynb
This Jupyter notebook contains the code for training a BERT-based model to perform multi-label classification of toxic comments. The primary elements of this notebook are:
- Imports: Incorporating necessary libraries and modules.
- Drive Mounting: Connecting to Google Drive to store the model and checkpoints.
- Labels and dictionaries: Specifying labels and generating dictionaries for label-to-index and index-to-label mapping.
- Dataset loading: Retrieving the dataset from a CSV file and tokenizing it using a pre-trained BERT tokenizer.
- Preprocessing: Developing a function to preprocess the data by tokenizing text and encoding labels.
- Model definition: Establishing the BERT-based model for multi-label classification.
- Dataset encoding: Encoding the dataset using the preprocessing function.
- Model training: Assembling and training the model using the encoded dataset.
- Upload to Huggingface: Using Huggingface cli to push the model and tokenizer to the hub.
app.py
This Python script is utilized to launch the trained model as a web application with the help of the Streamlit library. The primary components of this script are:
- Imports: Incorporating necessary libraries and modules.
- Title: Exhibiting the title of the web application.
- Input text area: Generating a text area for users to enter text for toxicity classification.
- Model choice: Offering a selection for users to pick their preferred model for classification.
- Tokenizer and model loading: Retrieving the tokenizer and model based on user selection.
- Sentiment pipeline: Developing a sentiment analysis pipeline using the loaded tokenizer and model.
- Submit button: Introducing a button for submitting the input text for classification.
- Classification and result display: When the submit button is activated, the input text is classified using the sentiment pipeline, and the results are presented in a table.
Dependencies
- Python 3.6+
- TensorFlow 2.x
- transformers
- datasets
- streamlit
- pandas
