Team Ai
Modelpublic

ariyul/gender_prediction_model_from_text

sourceHugging Facemitupdated 1y agoView on Hugging Face
2likes509downloads
Model Card

Gender Prediction from Text ✍️ β†’ πŸ‘©β€πŸ¦°πŸ‘¨

This model predicts the likely gender of an anonymous speaker or writer based solely on the content of an English text. It is built upon DeBERTa-v3-large and fine-tuned on a diverse, multilingual, and multi-domain dataset with both formal and informal texts.

πŸ“ Space link: πŸ”— Try it out on Hugging Face Spaces πŸ“ Model repo: πŸ”— View on Hugging Face Hub 🧠 Source code: GitHub


πŸ“Š Model Summary

  • β€”Base model: microsoft/deberta-v3-large
  • β€”Fine-tuned on: binary gender classification task (female vs male)
  • β€”Best F1 Score: 0.69 on a balanced multi-domain test set
  • β€”Max token length: 128
  • β€”Evaluation Metrics:
  • β€”F1: 0.69
  • β€”Accuracy: 0.69
  • β€”Precision: 0.69
  • β€”Recall: 0.69

πŸ“‚ Evaluation: View on Notebook


🧾 Datasets Used

DatasetDomainType
samzirbo/europarl.en-es.genderedFormal speech (Parliament)English
czyzi0/luna-speech-datasetPhone conversationsPolish β†’ Translated
czyzi0/pwr-azon-speech-datasetPhone conversationsPolish β†’ Translated
sagteam/author_profilingSocial postsRussian β†’ Translated
kaushalgawri/nptel-en-tags-and-gender-v0Spoken transcriptsEnglish
Blog Authorship CorpusBlog postsEnglish

All datasets were normalized, translated if necessary, deduplicated, and balanced via random undersampling to ensure equal representation of both genders.


πŸ› οΈ Preprocessing & Training

  • β€”Normalization: Cleaned quotes, dashes, placeholders, noise, and HTML/code from all datasets.
  • β€”Translation: Used Helsinki-NLP/opus-mt-* models for Polish and Russian data.
  • β€”Undersampling: Random undersampling to balance male and female samples.
  • β€”Training Strategy:
  • β€”LR Finder used to optimize learning rate (2.66e-6)
  • β€”Fine-tuned using early stopping on both F1 and loss
  • β€”Step-based evaluation every 250 steps
  • β€”Best checkpoint at step 24,750 saved and evaluated
  • β€”Second Phase Fine-tuning:
  • β€”Performed on full merged dataset for 2 epochs
  • β€”Used cosine learning rate scheduler and warm-up steps

πŸ“ˆ Performance (on full merged test set)

ClassPrecisionRecallF1-ScoreAccuracySupport
Female0.700.650.68591,027
Male0.680.720.70591,027
Macro Avg0.690.690.691,182,054
Accuracy0.691,182,054

πŸ“¦ Usage Example

python
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
import torch.nn.functional as F

device = torch.device("cuda" if torch.cuda.is_available() else "cpu")

model_name = "fc63/gender_prediction_model_from_text"
tokenizer = AutoTokenizer.from_pretrained(model_name, use_fast=False)
model = AutoModelForSequenceClassification.from_pretrained(model_name).eval().to(device)

def predict(text):
    inputs = tokenizer(text, return_tensors="pt", truncation=True, padding=True, max_length=128).to(device)
    with torch.no_grad():
        outputs = model(**inputs)
        probs = F.softmax(outputs.logits, dim=1)
    pred = torch.argmax(probs, dim=1).item()
    confidence = round(probs[0][pred].item() * 100, 1)
    gender = "Female" if pred == 0 else "Male"
    return f"{gender} (Confidence: {confidence}%)"
sample_text = "I love writing in my journal every night. It helps me reflect on the day and plan for tomorrow."
print(predict(sample_text))

The Output Of This Sample:

Female (Confidence: 84.1%)

πŸ“Œ Future Work & Limitations

I do not want to leave this model at the level of 0.69 accuracy and F1 score.

As far as I can detect at this point, there is a bias towards predicting emotional, psychological, and introspective texts as female. Similarly, more direct and result-oriented writings are also often predicted as male. Therefore, a large, carefully labeled dataset that reflects the opposite of this pattern is needed.

The datasets used to train this model had to be obtained from open-source platforms, which limited the range of accessible data.

To make further progress, I need to create and label a larger dataset myself β€” which requires a significant amount of time, effort, and cost.

Before moving to dataset creation, I plan to try a few more approaches using the current dataset. So far, alternative techniques have not helped improve the scores without causing overfitting. After testing a few more methods, if none work, the only step left will be building a new dataset β€” and that will likely be the point where I stop development, as it will be both labor-intensive and costly for me.


πŸ‘¨β€πŸ”¬ Author & License

Author: Furkan Γ‡oban Project: CENG-481 Gender Prediction Model License: MIT