Team Ai
Modelpublic

DunnBC22/distilbert-base-multilingual-cased-language_detection

sourceHugging Faceapache-2.0updated 6mo agoView on Hugging Face
2likes17downloads
Model Card

distilbert-base-multilingual-cased-language_detection

This model is a fine-tuned version of distilbert-base-multilingual-cased on the None dataset. It achieves the following results on the evaluation set:

  • —Loss: 0.0595
  • —Accuracy: 0.9971
  • —F1
  • —Weighted: 0.9971
  • —Micro: 0.9971
  • —Macro: 0.9977
  • —Recall
  • —Weighted: 0.9971
  • —Micro: 0.9971
  • —Macro: 0.9974
  • —Precision
  • —Weighted: 0.9971
  • —Micro: 0.9971
  • —Macro: 0.9981

Model description

This is a classification model of 16 different languages.

For more information on how it was created, check out the following link: https://github.com/DunnBC22/NLPProjects/blob/main/Language%20Detection/Language%20Detection-%2010k%20Samples/languagedetection-10k.ipynb

Intended uses & limitations

This model is intended to demonstrate my ability to solve a complex problem using technology.

Training and evaluation data

Dataset Source: https://www.kaggle.com/datasets/basilb2s/language-detection

Input Word Length:

Length of Input Text (in Words)

Input Word Length By Class:

Length of Input Text (in Words) By Class

Class Distribution:

Class Distribution

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • —learning_rate: 2e-05
  • —trainbatchsize: 64
  • —evalbatchsize: 64
  • —seed: 42
  • —optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
  • —lrschedulertype: linear
  • —num_epochs: 3

Training results

Training LossEpochStepValidation LossAccuracyWeighted F1Micro F1Macro F1Weighted RecallMicro RecallMacro RecallWeighted PrecisionMicro PrecisionMacro Precision
1.07831.01280.15440.98230.98190.98230.98060.98230.98230.97980.98470.98230.9852
0.11892.02560.05950.99710.99710.99710.99770.99710.99710.99740.99710.99710.9981
0.06513.03840.04730.99710.99710.99710.99770.99710.99710.99740.99710.99710.9981

Framework versions

  • —Transformers 4.26.1
  • —Pytorch 1.12.1
  • —Datasets 2.9.0
  • —Tokenizers 0.12.1

License Notice

This model is a fine-tuned derivative of a pretrained model. Users must comply with the original model license.

Dataset Notice

This model was fine-tuned on third-party datasets which may have separate licenses or usage restrictions.