Team Ai
Modelpublic

Silly-Machine/TuPy-Bert-Base-Multilabel

sourceHugging Facemitupdated 3y agoView on Hugging Face
1likes32downloads
README.md92 linesDownload Raw Back to root
1---2license: mit3datasets:4- Silly-Machine/TuPyE-Dataset5language:6- pt7 8pipeline_tag: text-classification9base_model: neuralmind/bert-base-portuguese-cased10widget:11- text: 'Bom dia, flor do dia!!'12 13model-index:14  - name: Yi-34B15    results:16      - task:17          type: text-classfication18        dataset:19          name: TuPyE-Dataset20          type: Silly-Machine/TuPyE-Dataset21        metrics:22          - type: f123            value: 0.8424            name: F1-score25            verified: true26          - type: precision27            value: 0.8528            name: Precision29            verified: true30          - type: recall31            value: 0.8432            name: Recall33            verified: true 34---35 36## Introduction37 38 39TuPy-Bert-Base-Multilabel is a fine-tuned BERT model designed specifically for multilabel classification of hate speech in Portuguese. 40Derived from the [BERTimbau base](https://huggingface.co/neuralmind/bert-base-portuguese-cased), 41TuPy-Bert-Base-Multilabel is a refined solution for addressing categorical hate speech concerns (ageism, aporophobia, body shame, capacitism, LGBTphobia, political, 42racism, religious intolerance, misogyny, and xenophobia).43For more details or specific inquiries, please refer to the [BERTimbau repository](https://github.com/neuralmind-ai/portuguese-bert/).44 45The efficacy of Language Models can exhibit notable variations when confronted with a shift in domain between training and test data. 46In the creation of a specialized Portuguese Language Model tailored for hate speech classification,47the original BERTimbau model underwent fine-tuning processe carried out on 48the [TuPy Hate Speech DataSet](https://huggingface.co/datasets/Silly-Machine/TuPyE-Dataset), sourced from diverse social networks.49 50## Available models51 52| Model                                    | Arch.      | #Layers | #Params |53| ---------------------------------------- | ---------- | ------- | ------- |54| `Silly-Machine/TuPy-Bert-Base-Binary-Classifier`  | BERT-Base	|12	|109M|55| `Silly-Machine/TuPy-Bert-Large-Binary-Classifier` | BERT-Large | 24      | 334M    |56| `Silly-Machine/TuPy-Bert-Base-Multilabel` | BERT-Base | 12      | 109M    |57| `Silly-Machine/TuPy-Bert-Large-Multilabel` | BERT-Large | 24      | 334M    |58 59## Example usage60 61```python62from transformers import AutoModelForSequenceClassification, AutoTokenizer, AutoConfig63import torch64import numpy as np65from scipy.special import softmax66 67def classify_hate_speech(model_name, text):68    model = AutoModelForSequenceClassification.from_pretrained(model_name)69    tokenizer = AutoTokenizer.from_pretrained(model_name)70    config = AutoConfig.from_pretrained(model_name)71 72    # Tokenize input text and prepare model input73    model_input = tokenizer(text, padding=True, return_tensors="pt")74 75    # Get model output scores76    with torch.no_grad():77        output = model(**model_input)78        scores = softmax(output.logits.numpy(), axis=1)79        ranking = np.argsort(scores[0])[::-1]80 81    # Print the results82    for i, rank in enumerate(ranking):83        label = config.id2label[rank]84        score = scores[0, rank]85        print(f"{i + 1}) Label: {label} Score: {score:.4f}")86 87# Example usage88model_name = "Silly-Machine/TuPy-Bert-Base-Multilabel"89text = "Bom dia, flor do dia!!"90classify_hate_speech(model_name, text)91 92```