Team Ai
Modelpublic

gamusa/vietnamese-toxic-comment-classification

sourceHugging Faceupdated 4mo agoView on Hugging Face
0likes
Model Card

DLNLP - Vietnamese Hate Speech Detection

This repository contains a local notebook workflow for Vietnamese hate speech detection.

Project Structure

  • —data/raw: raw CSV files
  • —data/processed: cleaned CSV files generated by preprocessing
  • —notebooks: notebook pipeline
  • —outputs/figures: saved charts and confusion matrices
  • —outputs/models: trained model artifacts
  • —outputs/results: metrics, reports, and prediction files

Dữ liệu đầu vào

The preprocessing notebook expects these files in data/raw:

  • —train_raw.csv
  • —val_raw.csv
  • —test_raw.csv

Each file should contain at least:

  • —free_text
  • —label_id

Setup

Create and activate a virtual environment, then install dependencies:

bash
python -m venv .venv
.venv\Scripts\activate
pip install -r requirements.txt

If you already have a working Python environment, installing from requirements.txt is enough.

How to run locally

Run the notebooks in this order:

  1. 1.notebooks/01_data_preprocessing.ipynb
  2. 2.notebooks/02_baseline_tfidf_svm.ipynb
  3. 3.notebooks/03_bilstm_train_evaluate.ipynb
  4. 4.notebooks/04_phobert_train_evaluate.ipynb
  5. 5.notebooks/05_results_error_analysis.ipynb

Demo giao diện

Sau khi đã có các file model trong outputs/models, chạy giao diện demo bằng:

bash
streamlit run app.py

Notes

  • —To make the repository clone-and-run for the demo app, keep the trained artifacts in outputs/models.
  • —Do not commit local environment folders such as .venv, notebook caches, or temporary training checkpoints.
  • —The notebooks resolve the project root automatically from the local workspace, so Colab Drive mounting is not required.
  • —Notebook 1 creates data/processed from the raw CSV files.
  • —Notebooks 2, 3, and 4 write outputs into outputs/models, outputs/results, and outputs/figures.
  • —Notebook 5 reads the artifacts from those local output folders and performs comparison/error analysis.
  • —For notebook 3 and notebook 4, the package list includes pyvi, transformers, torch, and accelerate.

Expected outputs

After a full run, you should see files such as:

  • —data/processed/train_processed.csv
  • —outputs/models/svm_model.joblib
  • —outputs/models/bilstm_best.pt
  • —outputs/models/phobert_base/
  • —outputs/results/svm_metrics.csv
  • —outputs/results/bilstm_metrics.csv
  • —outputs/results/phobert_metrics.csv
  • —outputs/results/phobert_test_predictions.csv
  • —outputs/figures/*.png