Silly-Machine/TuPy-Bert-Base-Multilabel
132
1---2license: mit3datasets:4- Silly-Machine/TuPyE-Dataset5language:6- pt7 8pipeline_tag: text-classification9base_model: neuralmind/bert-base-portuguese-cased10widget:11- text: 'Bom dia, flor do dia!!'12 13model-index:14 - name: Yi-34B15 results:16 - task:17 type: text-classfication18 dataset:19 name: TuPyE-Dataset20 type: Silly-Machine/TuPyE-Dataset21 metrics:22 - type: f123 value: 0.8424 name: F1-score25 verified: true26 - type: precision27 value: 0.8528 name: Precision29 verified: true30 - type: recall31 value: 0.8432 name: Recall33 verified: true 34---35 36## Introduction37 38 39TuPy-Bert-Base-Multilabel is a fine-tuned BERT model designed specifically for multilabel classification of hate speech in Portuguese. 40Derived from the [BERTimbau base](https://huggingface.co/neuralmind/bert-base-portuguese-cased), 41TuPy-Bert-Base-Multilabel is a refined solution for addressing categorical hate speech concerns (ageism, aporophobia, body shame, capacitism, LGBTphobia, political, 42racism, religious intolerance, misogyny, and xenophobia).43For more details or specific inquiries, please refer to the [BERTimbau repository](https://github.com/neuralmind-ai/portuguese-bert/).44 45The efficacy of Language Models can exhibit notable variations when confronted with a shift in domain between training and test data. 46In the creation of a specialized Portuguese Language Model tailored for hate speech classification,47the original BERTimbau model underwent fine-tuning processe carried out on 48the [TuPy Hate Speech DataSet](https://huggingface.co/datasets/Silly-Machine/TuPyE-Dataset), sourced from diverse social networks.49 50## Available models51 52| Model | Arch. | #Layers | #Params |53| ---------------------------------------- | ---------- | ------- | ------- |54| `Silly-Machine/TuPy-Bert-Base-Binary-Classifier` | BERT-Base |12 |109M|55| `Silly-Machine/TuPy-Bert-Large-Binary-Classifier` | BERT-Large | 24 | 334M |56| `Silly-Machine/TuPy-Bert-Base-Multilabel` | BERT-Base | 12 | 109M |57| `Silly-Machine/TuPy-Bert-Large-Multilabel` | BERT-Large | 24 | 334M |58 59## Example usage60 61```python62from transformers import AutoModelForSequenceClassification, AutoTokenizer, AutoConfig63import torch64import numpy as np65from scipy.special import softmax66 67def classify_hate_speech(model_name, text):68 model = AutoModelForSequenceClassification.from_pretrained(model_name)69 tokenizer = AutoTokenizer.from_pretrained(model_name)70 config = AutoConfig.from_pretrained(model_name)71 72 # Tokenize input text and prepare model input73 model_input = tokenizer(text, padding=True, return_tensors="pt")74 75 # Get model output scores76 with torch.no_grad():77 output = model(**model_input)78 scores = softmax(output.logits.numpy(), axis=1)79 ranking = np.argsort(scores[0])[::-1]80 81 # Print the results82 for i, rank in enumerate(ranking):83 label = config.id2label[rank]84 score = scores[0, rank]85 print(f"{i + 1}) Label: {label} Score: {score:.4f}")86 87# Example usage88model_name = "Silly-Machine/TuPy-Bert-Base-Multilabel"89text = "Bom dia, flor do dia!!"90classify_hate_speech(model_name, text)91 92```